DeepSeek V4 Pro 0813 Meluncur: Skor Coding Mendekati Model Termahal Dunia
DeepSeek merilis versi final model flagship-nya secara diam-diam β tanpa blog post, tanpa pengumuman besar. Tandanya hanya satu: tabel di halaman pricing API yang kini menampilkan versi baru DeepSeek-V4-Pro-0813 menggantikan versi preview yang bertahan sejak April.
Meski rilisnya senyap, klaimnya tidak kecil. DeepSeek menyebut model ini hanya kalah sekitar 5% dari Claude Fable 5 di sembilan benchmark agen, dengan harga yang jauh lebih murah. Bahkan, di dua benchmark DeepSeek justru menang.
Apa Itu DeepSeek V4 Pro 0813?
DeepSeek V4 Pro 0813 adalah build general availability (GA) dari DeepSeek V4 Pro β model yang sebelumnya berstatus preview sejak 24 April 2026. Rilis resminya tercatat pada 12 Agustus 2026, bersamaan dengan peluncuran Grok 4.6 dari SpaceXAI.
Spesifikasi intinya:
- Arsitektur: Mixture-of-Experts (MoE), 1,6 triliun parameter total, 49 miliar aktif per token
- Context window: 1 juta token, output maksimal 384 ribu token
- Mode: non-thinking, thinking, dan reasoning effort low/high/max
- Fitur agen: tool calling, JSON output, Responses API, kompatibel API Anthropic, FIM completion (mode non-thinking)
- Lisensi: MIT (untuk weights preview; weights versi 0813 belum dirilis)
- Harga API: $0,435 per juta token input (cache miss), $0,0036 cache hit, $0,87 per juta token output
Model ini memakai arsitektur hybrid attention (Compressed Sparse Attention + Heavily Compressed Attention) yang diklaim memangkas biaya komputasi per token hingga 73% dan ukuran KV cache hingga 90% dibanding generasi V3.2 pada konteks satu juta token. Kedua model V4 dilatih di atas 32 triliun token.
Benchmark Resmi: Klaim DeepSeek
DeepSeek mempublikasikan angka benchmark resminya dari model card dan harness internal (belum ada evaluator independen yang mereplikasi untuk build 0813). Pada effort maksimal (V4-Pro-Max), angka yang dilaporkan:
| Benchmark | Skor | Catatan |
|---|---|---|
| LiveCodeBench | 93,5 | Tertinggi di tabel vendor, di atas Gemini-3.1-Pro (91,7) dan Opus 4.6 (88,8) |
| SWE-bench Verified | 80,6 | Setara Gemini-3.1-Pro, tipis di bawah Opus 4.6 (80,8) |
| Terminal-Bench 2.0 | 67,9 | Di bawah GPT-5.4 xHigh (75,1) |
| GPQA Diamond | 90,1 | β |
| Humanityβs Last Exam | 37,7 | Di bawah Gemini-3.1-Pro (44,4) |
| MMLU-Pro | 87,5 | β |
| Codeforces rating | 3.206 | β |
Angka-angka ini untuk versi preview. Untuk build GA 0813, DeepSeek menerbitkan tabel perbandingan yang beredar lewat grup WeChat resminya, membandingkan langsung dengan kompetitor:
| Benchmark | DS V4 Pro 0813 | Claude Fable 5 | GLM-5.2 | Opus 4.8 |
|---|---|---|---|---|
| Terminal-Bench 2.1 | 87,9 | 88,0 | 81,0 | 85,0 |
| DeepSWE | 62,7 | 70,0 | 46,2 | 58,0 |
| HLE dengan tools | 60,0 | 63,0 | β | β |
| Cybergym | 83,3 | 83,1 | β | β |
| Toolathlon-Verified | 74,1 | 77,9 | β | β |
| AutomationBench | 31,8 | 29,1 | 12,9 | 27,2 |
| DSBench-FullStack | 71,1 | 77,2 | 61,8 | 71,6 |
Rata-rata keunggulan Fable 5 atas DeepSeek di sembilan benchmark: 5,3%. Jika Humanityβs Last Exam tanpa tools dikeluarkan (selisih 10,6 poin di sana paling timpang), rata-rata menyusut menjadi 2,8%. Angka ini dilaporkan vendor sendiri, jadi wajar bila dibaca dengan hati-hati.
Benchmark Independen: Artificial Analysis
Artificial Analysis, evaluator pihak ketiga, sudah memasukkan DeepSeek V4 Pro 0813 ke Intelligence Index v4.1.1 (gabungan 9 evaluasi termasuk Terminal-Bench v2.1, GDPval-AA, dan GPQA Diamond):
| Model | AA Index | Peringkat |
|---|---|---|
| Claude Opus 5 (max) | 63 | #1 dari 184 |
| GPT-5.6 Sol (max) | 61 | β |
| Grok 4.6 (high) | 61 | β |
| Kimi K3 (max) | 60 | #1 open weights (103) |
| Qwen3.8 Max | 58 | #10 dari 184 |
| GPT-5.6 Terra (max) | 57 | β |
| DeepSeek V4 Pro 0813 (max) | 53 | #2 open weights (104) |
| GLM-5.2 (max) | 53 | β |
| DeepSeek V4 Flash 0731 | 52 | β |
DeepSeek V4 Pro 0813 berada di urutan kedua model open weights, kalah 10 poin dari Opus 5 dan 8 poin dari GPT-5.6 Sol. Biaya evaluasi per tugas di AA tercatat $0,06 β jauh di bawah Opus 5 ($2,34) dan Sol ($1,23).
Perhitungan komunitas di Hacker News (geometric mean 7 benchmark coding/agen): GPT-5.6 Sol 65,5 > Fable 5 64,5 > Opus 5 64,0 > DeepSeek V4 Pro 0813 62,5 > Kimi K3 62,3 > DeepSeek V4 Flash 55,8 > GLM-5.2 47,3.
Perbandingan Harga: Di Mana DeepSeek Menang Telak
| Model | Input/M token | Output/M token | Cache hit |
|---|---|---|---|
| DeepSeek V4 Pro 0813 | $0,435 | $0,87 | $0,0036 |
| DeepSeek V4 Flash 0731 | $0,14 | $0,28 | $0,0028 |
| Claude Fable 5 | $10,00 | $50,00 | $1,00 |
| Claude Opus 5 | $5,00 | $25,00 | $0,50 |
| GPT-5.6 Sol | $5,00 | $30,00 | $0,50 |
| GPT-5.6 Luna | $0,20 | $1,20 | $0,02 |
| Kimi K3 | $3,00 | $15,00 | $0,30 |
| Qwen3.8 Max | $2,00 | $6,00 | $0,25 |
| GLM-5.2 | $1,40 | $4,40 | $0,235 |
| Grok 4.6 | $2,00 | $6,00 | $0,50 |
Dengan asumsi biaya campuran (blended) per tugas, Decrypt menghitung biaya Fable 5 sekitar $30 per tugas dibanding $0,65 untuk DeepSeek V4 Pro β sekitar 46 kali lebih mahal. Artificial Analysis mengukur biaya per tugas Fable 5 di $3,15 versus $0,03 untuk V4 Flash (105 kali). Keunggulan ini makin besar di beban kerja agen karena cache read DeepSeek sangat murah ($0,0036 per juta token), sementara pola penggunaan agen khasnya 82 ribu token cache hit per permintaan.
Ranking Per Kategori
Coding terbaik (mutlak): Claude Opus 5 β memimpin AA Index (63), dan Decrypt mencatat Opus 5 mengungguli Fable 5 di mayoritas benchmark dengan harga setengahnya. Fable 5 tetap juara di Terminal-Bench (88,0) dan DeepSWE (70,0).
Agen/terminal terbaik: Fable 5 (Terminal-Bench 2.1 88,0, DeepSWE 70,0), disusul DeepSeek V4 Pro 0813 yang hanya tertinggal 0,1 poin di Terminal-Bench dan menang di Cybergym serta AutomationBench.
Terbaik per dolar: DeepSeek V4 Pro 0813 β kelas Fable 5 dengan selisih 5%, harga 46 kali lebih murah, open weights MIT, dan cache termurah di pasar. Untuk volume ekstrem, V4 Flash 0731 (AA 52, $0,28 output) tetap juara ekonomi.
Open weights terbaik: Kimi K3 (AA 60) untuk kecerdasan umum; DeepSeek V4 Pro 0813 (AA 53) untuk nilai keseluruhan β dengan catatan lisensi Kimi K3 melarang penggunaan komersial, sementara DeepSeek MIT.
Posisi DeepSeek vs Opus 5 dan GPT-5.6 Sol
- vs Claude Opus 5: kalah 10 poin AA Index (53 vs 63) dan kalah di DeepSWE/SWE-bench, tapi harga output 28 kali lebih murah ($0,87 vs $25) dan open weights. Untuk tim yang butuh hasil maksimal tanpa batas budget: Opus 5. Untuk skala: DeepSeek.
- vs GPT-5.6 Sol: kalah 8 poin AA (53 vs 61), kalah di Terminal-Bench v3.0 dan DeepSWE (Sol Max 73%), tapi 34 kali lebih murah di output dan jauh lebih ringkas β Sol menghasilkan biaya evaluasi $2.823 di AA vs $135 untuk DeepSeek.
- Kesimpulan umum: DeepSeek V4 Pro 0813 menempati tier frontier kedua β di atas GLM-5.2 dan Qwen3.8 Max untuk coding, sejajar Kimi K3, sekitar 5-15% di bawah trio Opus 5/Sol/Fable dengan harga 20-46 kali lebih murah.
Catatan Penting
- Benchmark GA 0813 yang beredar adalah laporan vendor (DeepSeek Harness βakan dirilisβ), belum direplikasi independen.
- Weights versi 0813 belum dipublikasikan di Hugging Face β yang tersedia masih build preview April.
- DeepSeek mengumumkan akan menaikkan harga API secara signifikan dalam waktu dekat. Harga $0,435/$0,87 adalah harga saat ini, bukan harga permanen.
- Tes dunia nyata komunitas (HN) campur aduk: ada yang melaporkan Grok 4.6 menyelesaikan tugas tanpa bug dalam 3 menit dengan biaya $1,41, sementara DeepSeek butuh 12 menit dan $0,12 dengan bug β tetapi ini sampel satu percobaan, bukan data statistik.
Sumber
- DeepSeek API Docs β Models & Pricing β https://api-docs.deepseek.com/quick_start/pricing
- DeepSeek API Docs β Quick Start β https://api-docs.deepseek.com/
- Hugging Face β DeepSeek-V4-Flash-0731 model card (benchmark resmi) β https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731
- arXiv β DeepSeek-V4 Technical Report β https://arxiv.org/abs/2606.19348
- Artificial Analysis β DeepSeek V4 Pro 0813 β https://artificialanalysis.ai/models/deepseek-v4-pro
- Artificial Analysis β Leaderboard β https://artificialanalysis.ai/leaderboards/models
- OpenAI β API Pricing β https://platform.openai.com/docs/pricing
- Anthropic β Plans & Pricing β https://www.anthropic.com/pricing
- Decrypt β βChinaβs DeepSeek Upgrades V4 Pro: Claude Fable Is Only 5% Better at 4,500% the Priceβ β https://decrypt.co/375507/china-deepseek-upgrades-v4-pro-claude-fable
- Unite.AI β βDeepSeek Ships V4 Pro as Its Flagship Model Leaves Previewβ β https://www.unite.ai/deepseek-ships-v4-pro-as-its-flagship-model-leaves-preview/
- VentureBeat β βSpaceXAI debuts Grok 4.6β¦β β https://venturebeat.com/ai/spacexai-debuts-grok-4-6-overtaking-kimi-k3-performance/
- OpenRouter β DeepSeek V4 Pro 0813 β https://openrouter.ai/deepseek/deepseek-v4-pro-0813
- Hacker News β Thread DeepSeek V4 Pro 0813 β https://news.ycombinator.com/item?id=49274600