Grok-1 314B 在 M2 Ultra 上跑起来了
Georgi Gerganov 在推上说 Causally running Grok-1 at home(在家中随意运行 Grok-1)
从输出信息看,作者的电脑是Apple M2 Ultra
具体参数是:4-bit! iq3_s, 130gb metal buffer, 9 t/s
看来对我们来说并不那么随意,但至少不用8块H100了。
ggml_metal_init: allocatingggml_metal_init: found device: Apple M2 Ultraggml_metal_init: picking default device: Apple M2 Ultraggml_metal_init: default.metallib not found, loading from sourceggml_metal_init: GGML_METAL_PATH_RESOURCES = nilggml_metal_init: loading '/Users/ggerganov/development/github/llama.cpp/ggml-metal.metal'ggml_metal_init: GPU name: Apple M2 Ultraggml_metal_init: GPU family: MTLGPUFamilyApple8 (1008)ggml_metal_init: GPU family: MTLGPUFamilyCommon3 (3003)ggml_metal_init: GPU family: MTLGPUFamilyMetal3 (5001)ggml_metal_init: simdgroup reduction support = trueggml_metal_init: simdgroup matrix mul. support = trueggml_metal_init: hasUnifiedMemory = trueggml_metal_init: recommendedMaxWorkingSetSize = 154618.82 MBggml_backend_metal_buffer_type_alloc_buffer: allocated buffer, size = 128.00 MiB, (131511.92 / 147456.00)llama_kv_cache_init: Metal KV buffer size = 128.00 MiBllama_new_context_with_model: KV self size = 128.00 MiB, K (f16): 64.00 MiB, V (f16): 64.00 MiBllama_new_context_with_model: CPU output buffer size = 256.00 MiBggml_backend_metal_buffer_type_alloc_buffer: allocated buffer, size = 356.03 MiB, (131867.95 / 147456.00)llama_new_context_with_model: Metal compute buffer size = 356.03 MiBllama_new_context_with_model: CPU compute buffer size = 13.00 MiBllama_new_context_with_model: graph nodes = 3782llama_new_context_with_model: graph splits = 2system_info: n_threads = 16 / 24 | AVX = 0 | AVX_VNNI = 0 | AVX2 = 0 | AVX512 = 0 | AVX512_VBMI = 0 | AVX512_VNNI = 0 | FMA = 0 | NEON = 1 | ARM_FMA = 1 | F16C = 0 | FP16_VA = 1 | WASM_SIMD = 0 | BLAS = 1 | SSE3 = 0 | SSSE3 = 0 | VSX = 0 | MATMUL_INT8 = 0 |sampling:repeat_last_n = 64, repeat_penalty = 1.000, frequency_penalty = 0.000, presence_penalty = 0.000top_k = 40, tfs_z = 1.000, top_p = 0.950, min_p = 0.050, typical_p = 1.000, temp = 0.800mirostat = 0, mirostat_lr = 0.100, mirostat_ent = 5.000sampling order:CFG -> Penalties -> top_k -> tfs_z -> typical_p -> top_p -> min_p -> temperaturegenerate: n_ctx = 512, n_batch = 2048, n_predict = 64, n_keep = 0
有兴趣的看看视频吧
llama.cpp 的Add grok-1 support PR上也有详细的讨论,有兴趣的去看看吧
https://github.com/ggerganov/llama.cpp/pull/6204