📊 Benchmarks de Coding
Ranking de modelos de IA en los principales benchmarks de programación. Datos recopilados de fuentes públicas a agosto de 2026.
⚠️ Los benchmarks tienen limitaciones: pueden ser "harness-tuned", contaminados o no representar tu caso de uso específico. Úsalos como orientación, no como verdad absoluta. Verifica con evaluaciones independientes como LMSYS Chatbot Arena o Artificial Analysis.
🏆 Ranking SWE-bench Pro
El benchmark más relevante para evaluar capacidad de ingeniería de software real (resolución de issues en repos multi-archivo).
Tabla completa de datos
| Modelo | Empresa | Contexto | $/M entrada | SWE-bench Pro | HumanEval | Open-weight |
|---|---|---|---|---|---|---|
| Claude Opus 4.8 | Anthropic | 200K | $5 | 79.4% | — | — |
| GLM-5 | Zhipu | 200K | $1 | 77.8% | — | ✓ |
| GPT-5.5 | OpenAI | 400K | $5 | 75.2% | 96.2% | — |
| Gemini 3.1 Pro | 2M | $2 | 71.8% | 94.6% | — | |
| Kimi K2.6 | Moonshot | 262K | $0.95 | 71.6% | — | — |
| DeepSeek V4 Pro | DeepSeek | 128K | $0.435 | 68.5% | — | ✓ |
| Claude Sonnet 4.6 | Anthropic | 1M | $3 | — | — | — |
| Claude Haiku 4.5 | Anthropic | 200K | $1 | — | — | — |
| GPT-5.5 Pro | OpenAI | 400K | $30 | — | — | — |
| GPT-5.4 | OpenAI | 400K | $2.5 | — | — | — |
| GPT-5.3-Codex | OpenAI | 400K | $1.75 | — | — | — |
| GPT-4.1 nano | OpenAI | 1M | $0.1 | — | 0.9% | — |
| Gemini 3.5 Flash | 1M | $1.5 | — | — | — | |
| Gemini 3 Flash | 1M | $0.5 | — | — | — | |
| Gemini 3.1 Flash-Lite | 1M | $0.25 | — | — | — | |
| Phi-4 (familia) | Microsoft | 16K | $0.125 | — | 0.9% | ✓ |
| MAI-Code-1-Flash | Microsoft | 128K | — | — | — | — |
| Grok 3 | xAI | 131K | $3 | — | 0.9% | — |
| Grok 3 Mini | xAI | 131K | $0.3 | — | — | — |
| Llama 4 Maverick | Meta | 1M | $0.27 | — | 0.9% | ✓ |
| Mistral Large 3 | Mistral | 256K | $0.5 | — | — | ✓ |
| Mistral Small | Mistral | 128K | $0.2 | — | 0.8% | ✓ |
| DeepSeek V4 Flash | DeepSeek | 128K | $0.14 | — | — | — |
| DeepSeek V3 | DeepSeek | 128K | $0.27 | — | 0.9% | ✓ |
| Qwen 3.6 Plus | Alibaba | 1M | $0.5 | — | — | — |
| Qwen3.5 (open-weight) | Alibaba | 131K | $0.3 | — | — | ✓ |
| MiniMax M2.5 | MiniMax | 1M | $0.3 | — | — | — |
| MiMo-V2-Pro | Xiaomi | 256K | $0 | — | — | ✓ |
| ERNIE 5 | Baidu | 128K | $0.4 | — | — | — |
| Hunyuan Large | Tencent | 256K | $0.4 | — | — | ✓ |
| Step 3.5 Flash | StepFun | 128K | $0.1 | — | — | — |
| GPT-5.6 Sol (xhigh) | OpenAI | 1M | $5 | — | — | — |
| Claude Opus 5 (max) | Anthropic | 1M | $5 | — | — | — |
| GPT-5.6 Sol (max) | OpenAI | 1M | $5 | — | — | — |
| GPT-5.6 Sol (high) | OpenAI | 1M | $5 | — | — | — |
| Claude Opus 5 (xhigh) | Anthropic | 1M | $5 | — | — | — |
| GPT-5.6 Terra (max) | OpenAI | 1M | $2 | — | — | — |
| Claude Opus 5 (high) | Anthropic | 1M | $5 | — | — | — |
| Claude Fable 5 (with fallback) | Anthropic | 1M | $10 | — | — | — |
| GPT-5.6 Sol (medium) | OpenAI | 1M | $5 | — | — | — |
| Kimi K3 (max) | Moonshot | 1.048576M | $3 | — | — | ✓ |
| Claude Opus 5 (medium) | Anthropic | 1M | $5 | — | — | — |
| Grok 4.5 (high) | SpaceXAI | 500K | $2 | — | — | — |
| Kimi K3 (low) | Moonshot | 1.048576M | $3 | — | — | ✓ |
| Claude Sonnet 5 (max) | Anthropic | 1M | $2 | — | — | — |
| GPT-5.6 Luna (max) | OpenAI | 1M | $0.2 | — | — | — |
| Muse Spark 1.1 (xhigh) | Meta | 1.048576M | $1.25 | — | — | — |
| GPT-5.6 Terra (xhigh) | OpenAI | 1M | $2 | — | — | — |
| GPT-5.6 Sol (low) | OpenAI | 1M | $5 | — | — | — |
| Gemini 3.6 Flash | 1M | $1.5 | — | — | — | |
| Gemini 3.1 Pro Preview | 1M | $2 | — | — | — | |
| GLM-5.2 (max) | Zhipu | 1M | $1.4 | — | — | ✓ |
| GPT-5.6 Luna (xhigh) | OpenAI | 1M | $0.2 | — | — | — |
| GPT-5.6 Terra (high) | OpenAI | 1M | $2 | — | — | — |
| Claude Opus 5 (low) | Anthropic | 1M | $5 | — | — | — |
| Claude Sonnet 5 (Non-reasoning) | Anthropic | 1M | $2 | — | — | — |
| Qwen3.7 Max | Alibaba | 1M | $2.5 | — | — | — |
| GPT-5.6 Sol (Non-reasoning) | OpenAI | 1M | $5 | — | — | — |
| GPT-5.6 Terra (medium) | OpenAI | 1M | $2 | — | — | — |
| GPT-5.6 Luna (high) | OpenAI | 1M | $0.2 | — | — | — |
| Motif 3 (Beta) | MotifTechnologies | 262K | $0 | — | — | — |
| Kimi K2.7 Code | Moonshot | 256K | $0.95 | — | — | ✓ |
| MiMo-V2.5-Pro | Xiaomi | 1M | $0.435 | — | — | ✓ |
| KAT-Coder-Pro V2 | KwaiKAT | 256K | $0.3 | — | — | — |
| Nex-N2-Pro | NexAGI | 262K | $0.5 | — | — | ✓ |
| Hy3 | Tencent | 256K | $0.136 | — | — | ✓ |
| Agnes 2.5 Pro Alpha | SapiensAI | 1M | $0.45 | — | — | — |
| DeepSeek V4 Pro (high) | DeepSeek | 1M | $0.435 | — | — | ✓ |
| Muse Spark | Meta | 262K | $0 | — | — | — |
| MiniMax-M3 | MiniMax | 1M | $0.3 | — | — | ✓ |
| GPT-5.6 Terra (low) | OpenAI | 1M | $2 | — | — | — |
| MiMo-V2.5 | Xiaomi | 1M | $0.14 | — | — | ✓ |
| Qwen3.7 Plus | Alibaba | 1M | $0.4 | — | — | — |
| Qwen3.6 Plus | Alibaba | 1M | $0.5 | — | — | — |
| Qwen3.6 27B | Alibaba | 262K | $0.6 | — | — | ✓ |
| Inkling Small | ThinkingMachines | 1M | $0.3 | — | — | ✓ |
| JT-4.1 Flash 236B A21B | ChinaMobile | 256K | $0 | — | — | — |
| GPT-5.6 Terra (Non-reasoning) | OpenAI | 1M | $2 | — | — | — |
¿Qué mide cada benchmark?
SWE-bench Pro
Benchmark de referencia para ingeniería de software: el modelo debe resolver issues reales de GitHub en repos multi-archivo. Mide capacidad agéntica completa: leer código, planificar, implementar y verificar.
Ver fuente →HumanEval
Benchmark clásico de generación de código Python: 164 problemas de programación con tests unitarios. Mide la capacidad de escribir funciones correctas a partir de docstrings.
Ver fuente →LiveCodeBench
Benchmark dinámico con problemas de competiciones de programación recientes (post-entrenamiento). Evita contaminación de datos.
Ver fuente →💡 Cómo interpretar los benchmarks
Un modelo top en SWE-bench puede no ser el mejor para tu caso. Si haces mucho prototyping rápido, un modelo "flash" barato puede ser más productivo que el #1 del ranking.
Los vendors optimizan para benchmarks específicos. SWE-bench Pro y LiveCodeBench son más difíciles de gamear, pero no son inmunes. Valida siempre con tu propio código.
Los modelos se actualizan cada 2-3 meses. Un ranking de enero puede ser obsoleto en abril. Fuentes fiables: Artificial Analysis, LiveBench, SWE-bench.