Kimi K3 tops Arena's coding leaderboardAn open-weight model is now competitive with frontier closed models on coding benchmarks.
Anthropic and Blackstone bet on AI 'implementation'The next big AI money may be in deployment services, not model IP.
How Anthropic runs large-scale code migrations with Claude CodeA concrete playbook for using coding agents on real, large legacy codebases rather than toy demos.
Why agents that pass tests fail in productionPassing evals is not the same as being safe to ship - a warning for anyone gating agent releases on benchmark scores.
Ramp expands tracking of AI token spendToken costs are becoming a line item finance teams actively want visibility into.
1Password for Claude: credentialed agent access without leaking secretsSolves a core blocker for letting agents touch real systems: how to grant access without handing over raw credentials.