EyesInAI
Switchyard
Groundyard
Connect
Benchmarks
Models
Learn
News
Home
Switchyard
Groundyard
Connect
Which model for your task?
Leaderboard
Deep Suite
Coding Models
Bake-offs
Measured
Compare
Trends
Try Models
Catalog
Providers
By Origin
By Layer
Hardware
Open Source
License Tiers
What we’re building
Explainers
Automation & Agents
Cut Costs
Case Studies
Papers
Notes
Principles
Methodology
Jargon
Help
Contact
Latest
Governance & Access
Pricing & Billing
Data Centers & Power
Articles
Investigated & Declined
Trending models
Status
Archive
Benchmark your use case
Loading benchmark data…
EyesInAI
·
Loading live benchmark data
Home
Benchmark
Settings
Loading models…
EyesInAI
·
Loading models…
All models
anthropic
claude-opus-5
Last tested Jul 25, 2026
Overall pass
89%
Avg latency
4838 ms
Context
1000k
Tools
Yes
Input $/1M
$5.00
Output $/1M
$25.00
Tests run
18
Passed
16/18
AI summary
Hardware to run it
Hosted:
$5 in / $25 out
per 1M
Self-host:
API only
Proprietary / weights not released — hosted API only.
Test results
⚡
Ping
Latency & availability — single-word reply
2552 ms
🧮
Reasoning
Basic math reasoning — show work, give answer
2635 ms
{ }
JSON Output
Structured output compliance — valid JSON with required keys
2353 ms
💻
Code Gen
Python function generation with docstring
6054 ms
🚀
Throughput
Token generation speed — 500-token long-form response
10348 ms
🔍
Context Recall
Retrieval from in-context data — 20-item list Q&A
Compare with another model
See the leaderboard
2920 ms
🔧
Tool Use
Function/tool calling — get_weather invocation
fail
📝
Summarization
Real-world: distill a news article into 3 bullet points
3904 ms
🏷️
Classification
Real-world: sentiment + category from customer review
2449 ms
📄
Data Extraction
Real-world: pull structured fields from an invoice
7125 ms
📋
Instruction Follow
Real-world: multi-rule compliance (5 sentences, no "the", etc.)
7326 ms
✅
Format Compliance
IFEval-style: 4 bullets, keyword inclusion/exclusion, no preamble
fail
🪡
Long-Context Needle
Find a 6-digit code buried in ~3.5k tokens of filler text
3693 ms
🧠
Multi-Step Logic
BBH-lite: boolean expression + web of lies + object counting
3588 ms