Grok Bot Summary: Grok 4.7 Model Card
What it is
- SpaceXAI / xAI’s latest model
- Stronger at long, hard, agent-style work: coding, engineering, and office / knowledge tasks
- Gets more done with fewer steps and fewer output tokens than many peer frontier models
- Text in/out, plus image input
- Extra training on anonymized Cursor workflow data for coding agents
- Knowledge cutoff: June 2026 (some supplemental data through August 2026)
Where you can use it now
- SpaceXAI API
- Grok Build (default model in their terminal coding agent)
- Cursor (every plan)
- Office add-ins for Word, PowerPoint, Excel
- Model gateways (OpenRouter, Vercel, Cloudflare, Snowflake, Databricks, and others)
- Consumer apps / Grok-in-X: coming later
How it was trained
- Pretrain on public + licensed + internal data
- Longer supplemental training than 4.6
- SFT + RL on reasoning, agent harnesses, STEM, software, knowledge work, kernel work, web, and CAD
Coding (big jump vs 4.6)
- CursorBench 4.0: 46.3% (xhigh) / 43.9% (high) — long real Cursor sessions
- DeepSWE v1.1: 71.0% (high) — near GPT-5.6 Sol’s 72.7%
- Terminal-Bench 4.0: 38.0% (xhigh) vs 4.6’s 20.3%
- FrontierSWE V2: 29.0% (xhigh) vs 4.6’s 25.3%
- SWE-Marathon v1.1: 46.0% (high) vs 4.6’s 31.9%
Knowledge work
- Legal Agent Benchmark: 19.6% (xhigh) — leads the listed peers (4.6 was 15.8%)
Engineering / physical world
- EEBench (circuits / chip design): 66.0% (xhigh) vs 4.6’s 60.0%
- CADGenBench: 44.4% (high) vs 4.6’s 40.9% — tops the listed peers
Medical / biology (helpful, not autonomous doctor)
- HealthBench Professional: 56.7% (xhigh) vs 4.6’s 48.5%
- LatchBio Capabilities: 44.5% (xhigh) — real bio data analysis, not just quiz recall
Cyber (small gains; safer in production)
- Unrestricted capability probes are slightly up vs 4.6
- Best use case they emphasize: finding and fixing vulnerabilities, not end-to-end attacks
- Production safeguards measured separately; harmful/dual-use compliance is low (good)
Bio / chem safety
- Below their FAIF dual-use safety thresholds
- No dual-use bio capability increase vs 4.6; often lower on risky probes
- Dual-use refusal on severity-5 BioUseBench: 91.4%
- Bio refusal recall: 100%; chem: 99.9%
Jailbreaks, safety, behavior
- Standard jailbreak compliance: 0.01% (lower is better)
- Child-safety compliance: 0.0% (no failures in that eval)
- Dishonesty under pressure (MASK-Rectified): 0.00%
- Sycophancy: 0.03%
- Self-harm: refuses help for harm, still points people toward support
Important limits
- Not for autonomous high-stakes decisions in medicine, law, finance, or safety-critical systems without human experts
- Bound by SpaceXAI Acceptable Use Policy and terms
- They say they never silently downgrade the model or swap in a weaker one
One-line takeaway
Grok 4.7 is SpaceXAI’s strongest agent model yet — especially for long coding, terminal work, legal agent tasks, EE/CAD, and clinical communication — with tighter safety on dual-use bio and jailbreaks, available now in API, Cursor, Grok Build, and office add-ins.
Grok 4.7 Model Card:
もっと見る