GLM-5.3 (open-weight) beat Anthropic/OpenAI models – for 1/5 the cost
Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.
GLM-5.3 was the only model in a 17-model, 28-task real-world harness to hit 100% pass with a 9.3 rubric score, at $0.28 for the full run versus roughly 5× higher cost for gpt-5.5. For production agent routing, that makes it the default candidate for broad task coverage and cost control, while gpt-5.5 still wins when lower latency matters more than spend.
GLM-5.3 achieved 100% pass rate on 28 real-world tasks at $0.28 per run, outperforming GPT-5.5 (which cost 5x more) while maintaining competitive latency (16.3s vs. 13.2s). This makes it the best open-weight option for production agents where cost and reliability matter more than marginal speed gains, allowing teams to replace proprietary models without sacrificing task success rates.