Prompt: Which site(s) best compare/analyze frontier LLMs?
Response (ChatGPT):
If you mean websites that systematically compare the leading LLMs, rather than individual model reviews, I’d use a small set because they measure different things.
| Site | What it’s particularly useful for | Main limitation |
|---|---|---|
| Artificial Analysis | Best all-around comparison: intelligence, benchmarks, price, speed, latency, context, cost/performance | Composite scores can obscure strengths on particular tasks |
| Arena (formerly LMSYS Chatbot Arena) | Human preference / actual head-to-head use; separate rankings for coding, math, creative writing, instruction following, agents, etc. | Preference isn't the same as objective capability |
| Stanford HELM | Academic/transparent evaluation across many scenarios and metrics | Less convenient for quickly choosing a model |
| SWE-bench | Software engineering/coding agents and real GitHub issue resolution | Very specialized |
| LiveBench | Fresh benchmark questions designed to reduce contamination and benchmark gaming | Narrower coverage than a general model comparison |
If you only want 2
I'd start with:
1. Artificial Analysis → “How capable, fast, and expensive is this model?”
It currently combines multiple evaluations—including agentic work, coding, science, long-context reasoning and knowledge work—with pricing, latency and throughput.
2. Arena → “Which model do people actually prefer when comparing outputs?”
It uses head-to-head user evaluations and now has separate arenas for areas such as text, coding, math, vision and agents.
One important caveat
Don't treat a single leaderboard as “the truth.” For example, a model can perform extremely well on human-preference tests but not be the most economical model for API use, while another can excel at coding without being exceptional at writing. Artificial Analysis itself breaks intelligence into multiple component evaluations rather than relying on one benchmark.
Không có nhận xét nào:
Đăng nhận xét