PostgreSQL码农集散地

Grok-1 314B 在 M2 Ultra 上跑起来了

Georgi Gerganov 在推上说 Causally running Grok-1 at home(在家中随意运行 Grok-1)

Image

从输出信息看,作者的电脑是Apple M2 Ultra

具体参数是:4-bit! iq3_s, 130gb metal buffer, 9 t/s

看来对我们来说并不那么随意,但至少不用8块H100了。

ggml_metal_init: allocatingggml_metal_init: found device: Apple M2 Ultraggml_metal_init: picking default device: Apple M2 Ultraggml_metal_init: default.metallib not found, loading from sourceggml_metal_init: GGML_METAL_PATH_RESOURCES = nilggml_metal_init: loading '/Users/ggerganov/development/github/llama.cpp/ggml-metal.metal'ggml_metal_init: GPU name:   Apple M2 Ultraggml_metal_init: GPU family: MTLGPUFamilyApple8  (1008)ggml_metal_init: GPU family: MTLGPUFamilyCommon3 (3003)ggml_metal_init: GPU family: MTLGPUFamilyMetal3  (5001)ggml_metal_init: simdgroup reduction support   = trueggml_metal_init: simdgroup matrix mul. support = trueggml_metal_init: hasUnifiedMemory              = trueggml_metal_init: recommendedMaxWorkingSetSize  = 154618.82 MBggml_backend_metal_buffer_type_alloc_buffer: allocated buffer, size =   128.00 MiB, (131511.92 / 147456.00)llama_kv_cache_init:      Metal KV buffer size =   128.00 MiBllama_new_context_with_model: KV self size  =  128.00 MiB, K (f16):   64.00 MiB, V (f16):   64.00 MiBllama_new_context_with_model:        CPU  output buffer size =   256.00 MiBggml_backend_metal_buffer_type_alloc_buffer: allocated buffer, size =   356.03 MiB, (131867.95 / 147456.00)llama_new_context_with_model:      Metal compute buffer size =   356.03 MiBllama_new_context_with_model:        CPU compute buffer size =    13.00 MiBllama_new_context_with_model: graph nodes  = 3782llama_new_context_with_model: graph splits = 2
system_info: n_threads = 16 / 24 | AVX = 0 | AVX_VNNI = 0 | AVX2 = 0 | AVX512 = 0 | AVX512_VBMI = 0 | AVX512_VNNI = 0 | FMA = 0 | NEON = 1 | ARM_FMA = 1 | F16C = 0 | FP16_VA = 1 | WASM_SIMD = 0 | BLAS = 1 | SSE3 = 0 | SSSE3 = 0 | VSX = 0 | MATMUL_INT8 = 0 | sampling: repeat_last_n = 64, repeat_penalty = 1.000, frequency_penalty = 0.000, presence_penalty = 0.000 top_k = 40, tfs_z = 1.000, top_p = 0.950, min_p = 0.050, typical_p = 1.000, temp = 0.800 mirostat = 0, mirostat_lr = 0.100, mirostat_ent = 5.000sampling order: CFG -> Penalties -> top_k -> tfs_z -> typical_p -> top_p -> min_p -> temperature generate: n_ctx = 512, n_batch = 2048, n_predict = 64, n_keep = 0

有兴趣的看看视频吧

llama.cpp 的Add grok-1 support PR上也有详细的讨论,有兴趣的去看看吧

https://github.com/ggerganov/llama.cpp/pull/6204