{
  "video_id": "reddit_1vr0v1q",
  "channel_slug": "singularity",
  "channel_handle": "r/singularity",
  "title": "Qwen3.8 (27b) performs better than GPT-5.6-Terra (Max) for Agentic tasks",
  "url": "https://www.reddit.com/r/singularity/comments/1vr0v1q/qwen38_27b_performs_better_than_gpt56terra_max/",
  "external_url": null,
  "upload_date": "20260817",
  "published_at": "2026-08-17T18:40:41+00:00",
  "transcript": "https://preview.redd.it/k43scsdkbzjh1.png?width=624&format=png&auto=webp&s=bde429dce5283c28a83b7f66cea22d0469c751a9\n\nArtifical Analysis **Agentic Index**\n\n|Model (max reasoning effort)|Score|\n|:-|:-|\n|Qwen3.8 Max|58|\n|GPT-5.6-Sol|58|\n|Qwen3.8 (27b)|51|\n|GPT-5.6-Terra|50|\n\n\n\n--- Top Comments ---\n\n\n[26 upvotes] I can run this model on my M5 Max 64GB Macbook and it outperforms Terra on Max thinking.   \n  \nThe **Agentic Index** is more important than the **Intelligence Index** because I don't care about multiple choice questions, obscure knowledge, language translation, etc.  \n  \nStill waiting for **Coding Agent Index** results. \n\n[20 upvotes] I saw that Cerebras will host this model at around 2,000 tok/s. I wish I could run this model at more than 5 tok/s on my hardware.\n\n[7 upvotes] My real world experience is it's not actually very good at agentic stuff. \n\nIt's still very Qwen-like in bee-lining towards the primary goal while ignoring your direction. E.g. \"We are build XYZ app. First, I want you to design an architecture using established design patterns, appropriate architecture choices, and good engineering principles, and create the according mermaid files. Then I want you to use documentation-first TDD -- write documentation first, then write tests, then write code to pass the tests -- to implement the app.\"\n\nAbout 30% of the time it's going to bypass the specific architecting step in whole, and the majority of the rest of the time it's going to minimally satisfy that constraint, often writing little more than stubs. I'd say 20% of the time I get a good actual architecture design. When it does it, it's very capable, but it doesn't like to do it. When it passes the architecture step, it tends to write very \"direct to solution code\" with a mix of various trendy patterns like microservices half-heartedly adhered to, with a good mix of \"everything is React actually\" type thinking.\n\nIt almost never uses documentation-first TDD. It goes straight to coding, *the\n\n[4 upvotes] What an incredible model tbh, it’s single handledly going to make me buy either a dgx spark or m5 max with 64 or 128gb ram. \n\nI’m just resisting every-time, the best version in 4-8 months might be it. Basically gpt 5.5 mediums/high like output if I can relate. And these models finally don’t seem bechmaxxed. \n\nPerhaps there is one weird thing with agentic work, to benchmax the benchmark the model needs to DO stuff which is very agentic, previously it could be bench maxed, now it seems harder and harder in some sense unless the model self can perform agentic workflows. ",
  "transcript_chars": 2580,
  "ingested_at": "2026-08-18T01:30:15.051264+00:00",
  "source": "reddit",
  "yt_meta": {
    "score": 78,
    "upvote_ratio": 0.83,
    "num_comments": 50,
    "author": "UnknownEssence",
    "is_self": true
  }
}