AI Briefing
AI Briefing Publikasi Riset AI Indonesia

"The gap between narrow AI and AGI is enormous. We need breakthrough research to cross it safely."

β€” Yann LeCun Β· Chief AI Scientist at Meta

DeepSeek V4 Pro 0813 Meluncur: Skor Coding Mendekati Model Termahal Dunia

Oleh Eko Hadi Murwanto Β· Β·
DeepSeek V4 Pro 0813 Meluncur: Skor Coding Mendekati Model Termahal Dunia

DeepSeek merilis versi final model flagship-nya secara diam-diam β€” tanpa blog post, tanpa pengumuman besar. Tandanya hanya satu: tabel di halaman pricing API yang kini menampilkan versi baru DeepSeek-V4-Pro-0813 menggantikan versi preview yang bertahan sejak April.

Meski rilisnya senyap, klaimnya tidak kecil. DeepSeek menyebut model ini hanya kalah sekitar 5% dari Claude Fable 5 di sembilan benchmark agen, dengan harga yang jauh lebih murah. Bahkan, di dua benchmark DeepSeek justru menang.

Apa Itu DeepSeek V4 Pro 0813?

DeepSeek V4 Pro 0813 adalah build general availability (GA) dari DeepSeek V4 Pro β€” model yang sebelumnya berstatus preview sejak 24 April 2026. Rilis resminya tercatat pada 12 Agustus 2026, bersamaan dengan peluncuran Grok 4.6 dari SpaceXAI.

Spesifikasi intinya:

  • Arsitektur: Mixture-of-Experts (MoE), 1,6 triliun parameter total, 49 miliar aktif per token
  • Context window: 1 juta token, output maksimal 384 ribu token
  • Mode: non-thinking, thinking, dan reasoning effort low/high/max
  • Fitur agen: tool calling, JSON output, Responses API, kompatibel API Anthropic, FIM completion (mode non-thinking)
  • Lisensi: MIT (untuk weights preview; weights versi 0813 belum dirilis)
  • Harga API: $0,435 per juta token input (cache miss), $0,0036 cache hit, $0,87 per juta token output

Model ini memakai arsitektur hybrid attention (Compressed Sparse Attention + Heavily Compressed Attention) yang diklaim memangkas biaya komputasi per token hingga 73% dan ukuran KV cache hingga 90% dibanding generasi V3.2 pada konteks satu juta token. Kedua model V4 dilatih di atas 32 triliun token.

Benchmark Resmi: Klaim DeepSeek

DeepSeek mempublikasikan angka benchmark resminya dari model card dan harness internal (belum ada evaluator independen yang mereplikasi untuk build 0813). Pada effort maksimal (V4-Pro-Max), angka yang dilaporkan:

BenchmarkSkorCatatan
LiveCodeBench93,5Tertinggi di tabel vendor, di atas Gemini-3.1-Pro (91,7) dan Opus 4.6 (88,8)
SWE-bench Verified80,6Setara Gemini-3.1-Pro, tipis di bawah Opus 4.6 (80,8)
Terminal-Bench 2.067,9Di bawah GPT-5.4 xHigh (75,1)
GPQA Diamond90,1β€”
Humanity’s Last Exam37,7Di bawah Gemini-3.1-Pro (44,4)
MMLU-Pro87,5β€”
Codeforces rating3.206β€”

Angka-angka ini untuk versi preview. Untuk build GA 0813, DeepSeek menerbitkan tabel perbandingan yang beredar lewat grup WeChat resminya, membandingkan langsung dengan kompetitor:

BenchmarkDS V4 Pro 0813Claude Fable 5GLM-5.2Opus 4.8
Terminal-Bench 2.187,988,081,085,0
DeepSWE62,770,046,258,0
HLE dengan tools60,063,0β€”β€”
Cybergym83,383,1β€”β€”
Toolathlon-Verified74,177,9β€”β€”
AutomationBench31,829,112,927,2
DSBench-FullStack71,177,261,871,6

Rata-rata keunggulan Fable 5 atas DeepSeek di sembilan benchmark: 5,3%. Jika Humanity’s Last Exam tanpa tools dikeluarkan (selisih 10,6 poin di sana paling timpang), rata-rata menyusut menjadi 2,8%. Angka ini dilaporkan vendor sendiri, jadi wajar bila dibaca dengan hati-hati.

Benchmark Independen: Artificial Analysis

Artificial Analysis, evaluator pihak ketiga, sudah memasukkan DeepSeek V4 Pro 0813 ke Intelligence Index v4.1.1 (gabungan 9 evaluasi termasuk Terminal-Bench v2.1, GDPval-AA, dan GPQA Diamond):

ModelAA IndexPeringkat
Claude Opus 5 (max)63#1 dari 184
GPT-5.6 Sol (max)61β€”
Grok 4.6 (high)61β€”
Kimi K3 (max)60#1 open weights (103)
Qwen3.8 Max58#10 dari 184
GPT-5.6 Terra (max)57β€”
DeepSeek V4 Pro 0813 (max)53#2 open weights (104)
GLM-5.2 (max)53β€”
DeepSeek V4 Flash 073152β€”

DeepSeek V4 Pro 0813 berada di urutan kedua model open weights, kalah 10 poin dari Opus 5 dan 8 poin dari GPT-5.6 Sol. Biaya evaluasi per tugas di AA tercatat $0,06 β€” jauh di bawah Opus 5 ($2,34) dan Sol ($1,23).

Perhitungan komunitas di Hacker News (geometric mean 7 benchmark coding/agen): GPT-5.6 Sol 65,5 > Fable 5 64,5 > Opus 5 64,0 > DeepSeek V4 Pro 0813 62,5 > Kimi K3 62,3 > DeepSeek V4 Flash 55,8 > GLM-5.2 47,3.

Perbandingan Harga: Di Mana DeepSeek Menang Telak

ModelInput/M tokenOutput/M tokenCache hit
DeepSeek V4 Pro 0813$0,435$0,87$0,0036
DeepSeek V4 Flash 0731$0,14$0,28$0,0028
Claude Fable 5$10,00$50,00$1,00
Claude Opus 5$5,00$25,00$0,50
GPT-5.6 Sol$5,00$30,00$0,50
GPT-5.6 Luna$0,20$1,20$0,02
Kimi K3$3,00$15,00$0,30
Qwen3.8 Max$2,00$6,00$0,25
GLM-5.2$1,40$4,40$0,235
Grok 4.6$2,00$6,00$0,50

Dengan asumsi biaya campuran (blended) per tugas, Decrypt menghitung biaya Fable 5 sekitar $30 per tugas dibanding $0,65 untuk DeepSeek V4 Pro β€” sekitar 46 kali lebih mahal. Artificial Analysis mengukur biaya per tugas Fable 5 di $3,15 versus $0,03 untuk V4 Flash (105 kali). Keunggulan ini makin besar di beban kerja agen karena cache read DeepSeek sangat murah ($0,0036 per juta token), sementara pola penggunaan agen khasnya 82 ribu token cache hit per permintaan.

Ranking Per Kategori

Coding terbaik (mutlak): Claude Opus 5 β€” memimpin AA Index (63), dan Decrypt mencatat Opus 5 mengungguli Fable 5 di mayoritas benchmark dengan harga setengahnya. Fable 5 tetap juara di Terminal-Bench (88,0) dan DeepSWE (70,0).

Agen/terminal terbaik: Fable 5 (Terminal-Bench 2.1 88,0, DeepSWE 70,0), disusul DeepSeek V4 Pro 0813 yang hanya tertinggal 0,1 poin di Terminal-Bench dan menang di Cybergym serta AutomationBench.

Terbaik per dolar: DeepSeek V4 Pro 0813 β€” kelas Fable 5 dengan selisih 5%, harga 46 kali lebih murah, open weights MIT, dan cache termurah di pasar. Untuk volume ekstrem, V4 Flash 0731 (AA 52, $0,28 output) tetap juara ekonomi.

Open weights terbaik: Kimi K3 (AA 60) untuk kecerdasan umum; DeepSeek V4 Pro 0813 (AA 53) untuk nilai keseluruhan β€” dengan catatan lisensi Kimi K3 melarang penggunaan komersial, sementara DeepSeek MIT.

Posisi DeepSeek vs Opus 5 dan GPT-5.6 Sol

  • vs Claude Opus 5: kalah 10 poin AA Index (53 vs 63) dan kalah di DeepSWE/SWE-bench, tapi harga output 28 kali lebih murah ($0,87 vs $25) dan open weights. Untuk tim yang butuh hasil maksimal tanpa batas budget: Opus 5. Untuk skala: DeepSeek.
  • vs GPT-5.6 Sol: kalah 8 poin AA (53 vs 61), kalah di Terminal-Bench v3.0 dan DeepSWE (Sol Max 73%), tapi 34 kali lebih murah di output dan jauh lebih ringkas β€” Sol menghasilkan biaya evaluasi $2.823 di AA vs $135 untuk DeepSeek.
  • Kesimpulan umum: DeepSeek V4 Pro 0813 menempati tier frontier kedua β€” di atas GLM-5.2 dan Qwen3.8 Max untuk coding, sejajar Kimi K3, sekitar 5-15% di bawah trio Opus 5/Sol/Fable dengan harga 20-46 kali lebih murah.

Catatan Penting

  1. Benchmark GA 0813 yang beredar adalah laporan vendor (DeepSeek Harness β€œakan dirilis”), belum direplikasi independen.
  2. Weights versi 0813 belum dipublikasikan di Hugging Face β€” yang tersedia masih build preview April.
  3. DeepSeek mengumumkan akan menaikkan harga API secara signifikan dalam waktu dekat. Harga $0,435/$0,87 adalah harga saat ini, bukan harga permanen.
  4. Tes dunia nyata komunitas (HN) campur aduk: ada yang melaporkan Grok 4.6 menyelesaikan tugas tanpa bug dalam 3 menit dengan biaya $1,41, sementara DeepSeek butuh 12 menit dan $0,12 dengan bug β€” tetapi ini sampel satu percobaan, bukan data statistik.

Sumber

  1. DeepSeek API Docs β€” Models & Pricing β€” https://api-docs.deepseek.com/quick_start/pricing
  2. DeepSeek API Docs β€” Quick Start β€” https://api-docs.deepseek.com/
  3. Hugging Face β€” DeepSeek-V4-Flash-0731 model card (benchmark resmi) β€” https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731
  4. arXiv β€” DeepSeek-V4 Technical Report β€” https://arxiv.org/abs/2606.19348
  5. Artificial Analysis β€” DeepSeek V4 Pro 0813 β€” https://artificialanalysis.ai/models/deepseek-v4-pro
  6. Artificial Analysis β€” Leaderboard β€” https://artificialanalysis.ai/leaderboards/models
  7. OpenAI β€” API Pricing β€” https://platform.openai.com/docs/pricing
  8. Anthropic β€” Plans & Pricing β€” https://www.anthropic.com/pricing
  9. Decrypt β€” β€œChina’s DeepSeek Upgrades V4 Pro: Claude Fable Is Only 5% Better at 4,500% the Price” β€” https://decrypt.co/375507/china-deepseek-upgrades-v4-pro-claude-fable
  10. Unite.AI β€” β€œDeepSeek Ships V4 Pro as Its Flagship Model Leaves Preview” β€” https://www.unite.ai/deepseek-ships-v4-pro-as-its-flagship-model-leaves-preview/
  11. VentureBeat β€” β€œSpaceXAI debuts Grok 4.6…” β€” https://venturebeat.com/ai/spacexai-debuts-grok-4-6-overtaking-kimi-k3-performance/
  12. OpenRouter β€” DeepSeek V4 Pro 0813 β€” https://openrouter.ai/deepseek/deepseek-v4-pro-0813
  13. Hacker News β€” Thread DeepSeek V4 Pro 0813 β€” https://news.ycombinator.com/item?id=49274600
Bagikan artikel ini
Telegram WhatsApp Facebook X

Artikel Terkait

Cognition AI Bidik Valuasi $40 Miliar: Devin Lebih Bernilai dari Cursor?

Cognition AI Bidik Valuasi $40 Miliar: Devin Lebih Bernilai dari Cursor?

13 Agustus 2026
Lovable Valuasi $13,3 Miliar: Startup Swedia Guncang Pasar AI Coding dengan Pendanaan $400 Juta

Lovable Valuasi $13,3 Miliar: Startup Swedia Guncang Pasar AI Coding dengan Pendanaan $400 Juta

13 Agustus 2026
Alibaba Rilis Qwen 3.8-Max: Model AI 2,4 Triliun Parameter Saingan Terberat Model AI China

Alibaba Rilis Qwen 3.8-Max: Model AI 2,4 Triliun Parameter Saingan Terberat Model AI China

5 Agustus 2026
← Kembali ke Beranda