Abu Dhabi · Apache 2.0 · vollständig offen (Gewichte, Code, Trainingsdaten, Checkpoints, Logs) · 6 Modelle von 0,9 B bis 375 B · veröffentlicht 03.09.2026
| Modell | Typ | VRAM Q4_K_M | 6 GB VRAMz. B. NVIDIA A2000 | 12 GB VRAMz. B. RTX 3060 Ti | 16 GB VRAMz. B. RTX 4060 Ti | 24 GB VRAMz. B. RTX 3090 / 4090 | CPU · 16 GB RAMz. B. Ryzen 7 (AM4) |
|---|---|---|---|---|---|---|---|
| K2-Horizon-0.9B | Dense | ~0,7 GB | Sehr komfortabelViel Puffer, auch bei 128k Kontext | Sehr komfortabel— | Sehr komfortabel— | Sehr komfortabel— | Gut nutzbarAuch auf schwacher CPU-Hardware flott |
| K2-Horizon-3.7B | Dense | ~2,3 GB | KomfortabelPasst mit Puffer für moderaten Kontext | Sehr komfortabel— | Sehr komfortabel— | Sehr komfortabel— | Gut nutzbar— |
| K2-Horizon-7B | Dense | ~4,4 GB | Knapp möglichWenig Puffer bei langem Kontext | Komfortabel— | Sehr komfortabel— | Sehr komfortabel— | Gut nutzbar— |
| K2-Horizon-32B | Dense | ~20 GB | Nicht möglich— | Nicht möglich— | Nicht möglichModell selbst passt knapp nicht, 512k-KV-Cache erst recht nicht | Knapp möglichNur bei stark reduziertem Kontext (weit unter 512k) | Nicht möglichSprengt 16 GB RAM |
| K2-Horizon-MoVA-36B-A4B | MoE | ~23 GB | Nicht möglich— | Nicht möglich— | Nicht möglichKnapp zu klein für Gewichte + Kontext | Knapp möglichPasst bei moderatem Kontext, 512k sprengt den Puffer | Gut nutzbarNur ~4 B aktive Parameter → trotz 36 B Gesamtgröße CPU-tauglich |
| K2-Horizon-375B-A23B | MoE | ~235 GB (Schätzung aus Parameterzahl) |
Nicht möglich— | Nicht möglich— | Nicht möglich— | Nicht möglichBraucht Multi-GPU-Cluster (Datacenter-Klasse) | Nicht möglich— |
| Modell | Text | Reasoning-Parser | Tool Calling | MoE | Kontext | Aktive Parameter |
|---|---|---|---|---|---|---|
| K2-Horizon-0.9B | ✓ | ✓ | ✓ | ✗ | 128k | 1,08 B |
| K2-Horizon-3.7B | ✓ | ✓ | ✓ | ✗ | ≥128k* | 3,7 B |
| K2-Horizon-7B | ✓ | ✓ | ✓ | ✗ | ≥128k* | 7 B |
| K2-Horizon-32B | ✓ | ✓ | ✓ | ✗ | 512k | 32 B |
| K2-Horizon-MoVA-36B-A4B | ✓ | ✓ | ✓ | ✓ | 512k | ≈ 4 B |
| K2-Horizon-375B-A23B | ✓ | ✓ | ✓ | ✓ | 512k | ≈ 23 B |
k2_horizon) ist in Upstream-llama.cpp zum Zeitpunkt der Veröffentlichung noch nicht gemergt (PR ausstehend). Für GGUF-Inferenz wird aktuell der Fork MBZUAI-IFM/llama.cpp (Branch model/K2Horizon) benötigt. Die VRAM-Zahlen für 3.7B/7B/32B/36B-A4B/375B-A23B sind Schätzungen auf Basis der Gesamtparameterzahl (analog zur Berechnung bei anderen Familien in dieser Übersicht); nur für 0.9B liegt eine offiziell bestätigte Parameterzahl inkl. Embeddings vor. * Für 3.7B und 7B ist keine offiziell bestätigte Kontextlänge auffindbar.