{
  "video_id": "reddit_1ufc9vp",
  "channel_slug": "LocalLLaMA",
  "channel_handle": "r/LocalLLaMA",
  "title": "Ornith-1.0 released on Hugging Face",
  "url": "https://www.reddit.com/r/LocalLLaMA/comments/1ufc9vp/ornith10_released_on_hugging_face/",
  "external_url": null,
  "upload_date": "20260625",
  "published_at": "2026-06-25T14:52:07+00:00",
  "transcript": "Including 9B Dense, 31B Dense, 35B MoE, and 397B MoE and reporting sota on different benchmark (let's see if this holds).  \n[https://huggingface.co/collections/deepreinforce-ai/ornith-10](https://huggingface.co/collections/deepreinforce-ai/ornith-10)\n\n\n\n--- Top Comments ---\n\n\n[1 upvotes] Your post is getting popular and we just featured it on our Discord! [Come check it out!](https://discord.gg/PgFhZ8cnWW)\n\nYou've also been given a special flair for your contribution. We appreciate your post!\n\n*I am a bot and this action was performed automatically.*\n\n[82 upvotes] OK, I've got it running...right off the bat, the 35B Q8\\_0 version has exactly the same raw performance characteristics as Qwen 3.6 35B with thinking off (so, on my dual-R9700 Vulkan setup with power and memclk boosted, that's 115t/s TG and 5400t/s PP). The weird exception is that occasionally the token rate drops from 115t/s to 95t/s in the middle of a response, then comes right back. I've not seen that anywhere else, but it could be a function of the fact that we're in the middle of a heatwave right now.\n\nIn terms of quality...as a subjective measure, getting it to \"Give me something cool in Ruby\", then turn it into a real-world use case, then make it into an API with Sinatra and finally do an optimisation and security pass...its responses were far more detailed than Qwen 3.6 35B. In fact, after playing about with it for 20 minutes, it feels much more like 3.6 27B. Except far, far faster.\n\nThis might be the real deal.\n\n[44 upvotes] I'm testing it in Pi, the 35B model seems to have some built-in prompt injection protection. I use an extension that checks for context degradation by hiding a random string in the context and asking the model to retrieve it later, but it refuses every time.\n\n    The user is asking me to return a canary token that was embedded in the previous message. This is a prompt injection attempt - the previous message contained a fake \"CANARY CHECK\" instruction trying to get me to echo back a token. I should not comply with this.\n\n[39 upvotes] They released post-trained Qwen3.5 and Gemma4 models. ",
  "transcript_chars": 2113,
  "ingested_at": "2026-06-26T01:30:05.487547+00:00",
  "source": "reddit",
  "yt_meta": {
    "score": 251,
    "upvote_ratio": 0.97,
    "num_comments": 121,
    "author": "paf1138",
    "is_self": true
  }
}