{
  "video_id": "reddit_1uqew7g",
  "channel_slug": "singularity",
  "channel_handle": "r/singularity",
  "title": "Open-source models are closing the coding gap with GPT/Claude/Gemini ~1.5x faster than the frontier is advancing, and on decontaminated benchmarks a 27B model already beats Claude Opus 4.8 [live dashboard + analysis]",
  "url": "https://www.reddit.com/r/singularity/comments/1uqew7g/opensource_models_are_closing_the_coding_gap_with/",
  "external_url": null,
  "upload_date": "20260708",
  "published_at": "2026-07-08T01:44:23+00:00",
  "transcript": "Everyone argues about whether open-source AI is catching up to the closed labs. I got tired of vibes, so I built a live dashboard that plots open-weight vs closed models on the coding benchmarks that matter (SWE-bench Verified, SWE-rebench, BFCL tool-calling, LiveCodeBench) over time, then ran the actual statistics on the trend.\n\nWhat the data says:\n\n- **Open small models are the steepest line on the board.** The best model you can run on a single consumer GPU went from 20% on SWE-bench Verified (Dec 2024) to 77% (mid-2026). Fitting the running-best frontier of each group, open ≤35B improves ~+39 pts/yr vs ~+26 for the closed frontier. That is ~1.5x faster, and the difference is statistically significant (p≈0.0002).\n- **On the benchmark that can't be gamed, the gap is almost gone.** SWE-rebench pulls fresh GitHub issues every month, so nothing is memorized. There, a 27B open model (Qwen3.5-27B) scores 58.9, within ~4 points of the global #1 and above Claude Opus 4.8 (56.5), even though Opus posts 88.6 on the public benchmark. Most of the visible \"closed lead\" is contamination, not capability.\n- **I deliberately do not predict a crossover date.** Extrapolating where two near-parallel lines cross is statistically unstable (the 95% interval runs mid-2026 to past 2028). The direction and rate are solid; the calendar date is not, so I don't headline one, and you should be skeptical of anyone who does.\n\nThe one thing genuinely holding open models back is not raw intelligence, it is tool-call reliability. On BFCL v4 it is Anthropic 77.5 / Google 72.5 / open ≤35B 51.4, and that gap is not closing. It is a data problem: the closed labs train on billions of real agent trajectories from their own products (Claude Code, Codex), and there is no open equivalent. The writeup ends with a concrete pitch: build an open harness that collects anonymized tool-call traces plus success labels and pools them into a public dataset anyone can train on. That is a coordination problem, which open source is good at, unlike a frontier pretraining run.\n\nDashboard (live, refreshes daily): https://botlab.dev/open-source-llm-benchmarks/\nFull writeup with the stats and charts: https://botlab.dev/open-models-closed-ai-crossover-2026\n\nData comes from benchlm.ai, swe-rebench.com, and the Berkeley BFCL leaderboard. (Disclosure: my own project, free, no signup, no ads.)\n\n\n\n--- Top Comments ---\n\n\n[50 upvotes] Dude... If you seriously think 27B is superior to Opus you either never actually used or are delusional. It is not, not even close. You can argue its close to Sonnet 4.6, but I say it is still a little bit behind.\n\nThis is not to say its a bad model, quite the contrary, for that size it is a beast. But its not even close to Opus\n\n[14 upvotes] the trend is real but the benchmark framing hides the actual gap. a 27b beating opus on decontaminated swe-bench is one-shot coding, where open models genuinely caught up (glm-5 is around 77% on swe-bench verified now). the gap that's NOT closing as fast is agentic reliability, tool-calling over long horizons. the small model nails the isolated task then falls apart on turn 15 of a real agent loop. one-shot is basically solved, the loop isn't.\n\n[11 upvotes] Can I do anything with this with my 3090 single GPU?",
  "transcript_chars": 3271,
  "ingested_at": "2026-07-08T13:30:29.636980+00:00",
  "source": "reddit",
  "yt_meta": {
    "score": 70,
    "upvote_ratio": 0.78,
    "num_comments": 31,
    "author": "toadlyBroodle",
    "is_self": true
  }
}