Gemini 3.5 Flash โ First Independent SWE-bench Pro Score Published: 55.1%, arriving after the May 21 digest scan. This positions Gemini 3.5 Flash above Gemini 3.1 Pro (54.2%) and confirms the capability delta over the prior generation, but places it behind Claude Opus 4.7 (64.3%, current #1) and GPT-5.5 (58.6%, #2). The 9.2-point gap versus Opus 4.7 is the concrete benchmark signal for developers evaluating model selection on software engineering agent tasks at scale. Context: Gemini 3.5 Flash's self-reported agentic benchmarks (Terminal-Bench 76.2%, MCP Atlas 83.6%) favor multi-step reasoning over long-horizon software engineering; SWE-bench Pro is the harder, longer-horizon test. Combined with yesterday's digest: Flash-tier pricing ($1.50/$9) with solid agentic benchmarks but a measurable SWE-bench gap โ the routing decision for coding agents depends on which benchmark category your workload resembles.
LMArena: Third-party leaderboard changelog (arena.ai/blog/leaderboard-changelog) reports gemini-3.5-flash was added to Text and Code leaderboards on May 19, 2026. Stable Elo not yet confirmed in this scan โ the May 21 digest reported no LMArena entry as of that scan; watch next cycle for first stable rating.