Rendered at 12:32:20 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
jjcm 15 hours ago [-]
China has caught up is the main takeaway here. The SOTA models are so close that it's really hard to compare them intelligence wise - you have to get a feel for them yourself and what works for you.
What I'm really excited for is the 27B model. 3.6 is still the king of local, and if 3.8 makes the same improvements it could really legitimately make local viable as a default. I'd love to run a perpetual agent on 3.8 that's locally driven.
monster_truck 14 hours ago [-]
Something I don't think many have internalized is that China has been as good or better for quite a while now (long before anyone was pointing distillation fingers) and enough people have finally tried it for themselves that the understanding has reached critical mass and the careful narrative of american companies is collapsing.
When I finally put $15 into Deepseek and it beat the brakes off Codex 5.5 on multiple rather complex projects without any of the obnoxious mistakes, I was sick to my stomach with buyers remorse. I couldn't believe I ever felt like I was getting my moneys worth at $200/mo. I wouldn't even use OAI's models if they were free and unlimited at this point, I'll happily pay for what I already know works. No reset bingo, no cache errors, no annoying shitposters as a primary source of info. Oh, and I still had $10 of tokens left
And yes, 3.6 is excellent locally. The rest of this year is gonna be awesome
4d4m 9 hours ago [-]
+1 most people are too afraid to try something new. They've been roughly on par with their frontier offerings for 6 months if not more.
aliasxneo 13 hours ago [-]
And I had the opposite experience. It's a really interesting phenomenon that I can't really explain. My co-founder swears by Deepseek and yet just the other day we were conversing and he was telling me about some of the issues with the way the AI was behaving and trying to show off the cool workarounds he came up with to limit it. I was like, "Interesting, yeah, I've literally never had that problem."
I suspect that the models are genuinely close and that certain experiences get felt across providers but are inconsistent enough to convince people one is superior to the other. I for one have tried Deepseek on and off since my co-founder is fond of it and I've stopped trying now because I never have a good experience.
On reddit et al., people talk about LLM brands like their sports teams.
I think the first-party ecosystem moats they're all trying to build are exacerbating this tendency, as now people have a lot of learning time sunk in a company-specific option.
wanderlust123 13 hours ago [-]
That’s not surprising. I have been using Deepseek and it consistently produces excellent output given the right context howevwr. It depends on the task as it does have blindspots.
surgical_fire 4 hours ago [-]
That may be true.
I switched to DeepSeek entirely once I decided to put 10 bucks on it and I realized that it could do whatever I was throwing at Claude or ChatGPT prior to that.
I recommended it to one of my friends, and he was surprised DeepSeek could solve task that Claude got stuck at. I was surprised at it too.
I know others that tried and were less impressed too.
kadushka 13 hours ago [-]
Opus 5 and 5.6 Sol are definitely not smart enough to do my job. They require constant supervision. So why would I want to switch to even worse model? Even if it's just slightly worse?
Cookingboy 13 hours ago [-]
>So why would I want to switch to even worse model?
There would be no reason to if you are in the privileged position where cost isn't an issue.
For the rest of us something that's 95% as good for 20% the price is a hell of a value proposition.
FuckButtons 10 hours ago [-]
Pareto optimal dominant vs a human for the same task, not an unreasonable framing but that assumes that it can actually do the task, which the op was arguing it couldn’t at all. Which, I suppose you could model as the utility of task completion % as being non linear. I have heard many people argue that the nature of work is messy and complicated and many things they do could not easily be emulated or automated. I do wonder how many of those activities are actually something that are connected to a companies ability to generate revenue or are just the messy interactions between people.
kadushka 10 hours ago [-]
Cost is absolutely an issue here - my time is worth approximately $1000/day, so if a slightly worse model wastes one more hour of my time a day than the best model, it costs the company >$2k/mo. Fortunately my employer understands this well and encourages me to use the best models as much as I can.
surgical_fire 4 hours ago [-]
This reply must have cost dozens of dollars.
ThunderSizzle 45 minutes ago [-]
This must be the new linked in strat. What I learned about using the best model after talking to a homeless person.
esafak 12 hours ago [-]
That's why every benchmark should show the Pareto frontier against cost and latency.
matheusmoreira 10 hours ago [-]
> So why would I want to switch to even worse model? Even if it's just slightly worse?
Self-hosting is the biggest reason.
Systemerror7A69 8 hours ago [-]
If they already require your constant supervision the reason is money.
surgical_fire 4 hours ago [-]
> They require constant supervision
I think that may be part of it.
LLMs can be autonomous to an extent. All of them need steering - which is why I feel they are more a superpower the more I am an expert on the subject matter.
The more you want it to be autonomous, than yeah, you may benefit from using the very best the industry has to offer, however slightly better it is.
But if you are always in the loop anyway, you may want to try DeepSeek. You will get similar results for a fraction of the price.
cyanydeez 13 hours ago [-]
I think americans assume when they see a chinese or asian person working at an american business that they "escaped" china as opposed to just being rich enough to go to school abroad. and has little to no bearing on the amount of intelligent going around.
14u2c 10 hours ago [-]
If Americans see an Asian person working at a US business, they will assume that person is an American. They may even ask what state you are from. It's honestly one of the nice things about the place, you can belong even if you are not from there.
pessimizer 13 hours ago [-]
They've been continuously programmed with insane beliefs about China, which is less shocking when you understand what insane beliefs that they've had programmed into them about their neighbors. The world and your neighborhood are full of evil communists and Nazis who are trying to kill you all the time.
TacticalCoder 11 hours ago [-]
> The world and your neighborhood are full of evil communists and Nazis who are trying to kill you all the time.
Don't know about that but your neighbors in Iran in early january happened to be "nice people" who just followed the orders to slaughter 30 000 unarmed civilians.
We could talk about the, what 600 000 deaths, including many civilians, in the Ukraine/Russia war.
Or we could talk about the number of nice palestinians killed since the beginning of the war in Gaza. Or we could go a bit further and talk about the joy and celebration in Gaza after their heroes brought back 200 hostages after having slaughtered 1200 civilians.
You may be living in a place that you think shields you from those but I know the ideologies behind these acts.
The fallacy of gray is just that: it's not true that there's always a nice middle ground and that there's no evil ideology out there.
Something something about the price of liberty being eternal vigilance. For there are people abusing your blind trust.
peterashford 11 hours ago [-]
QED
tovlier 11 hours ago [-]
[dead]
swat535 12 hours ago [-]
I don't think it's all Americans however there is a portion of them who are not able to grasp the world outside of their borders.
I think it's mainly due to poor education many receive and a very controlled media that suppresses information.
It's shocking considering how much money they spend on education compared to other nations.
nimchimpsky 10 hours ago [-]
[dead]
sieabahlpark 12 hours ago [-]
[dead]
doginasuit 15 hours ago [-]
Another potential takeaway is that the models all gathering around the same point supports the idea that there is a ceiling to LLM capability.
AustinDev 15 hours ago [-]
They always all gather around the same spot then that spot moves every 6-9 months. I think the clustering is more likely evidence of distillation. I don't personally think distillation is a bad thing. If the LLM providers can distill all of human output into their models for 'free'. I don't think distilling a model from the output of those models is morally wrong.
rlupi 5 hours ago [-]
I think it's more likely to be the effect of synchronization of launches, and the fact that models that do not challenge SOTA in some way do not get launched (think Gemini Pro delays), launched quietly or do not get any attention.
conception 12 hours ago [-]
If you talk to the Chinese models, even super smart Qwen 3.8, you can tell they are distilled just from the verbal ticks they have. Gemini, ChatGPT and Claude do not sound alike. The Chinese models 100% sound like one of the 3, usually Claude. American models are load bearing for this LLM generation seam.
kayson 11 hours ago [-]
- load bearing -
AustinDev 11 hours ago [-]
seams, boundaries, envelopes, etc
10 hours ago [-]
FuckButtons 10 hours ago [-]
That’s definitely my impression of deepseek 0731 after a fair bit of use via ds4, it sounds like Claude.
ofjcihen 14 hours ago [-]
I gathered that the most recent advances haven’t been in capabilities of the model but more the way that it’s able to be employed (most recently agents).
4 hours ago [-]
bossyTeacher 7 hours ago [-]
> LLM providers can distill all of human output into their models for 'free'
Not sure what part of being charged guilty and paying a fine you see as "free".
13 hours ago [-]
michelsedgh 14 hours ago [-]
What an interesting take. One question, do you think stealing from a thief is morally okay? I'm just asking no judgement on my side.
AussieWog93 14 hours ago [-]
I'd say it's more "Downloading LimeWire Pro from LimeWire" than actual theft.
dullcrisp 11 hours ago [-]
Why isn’t it more like building a hardware store using lumber you purchased from a competing hardware store? Or founding a school using an education you obtained at a different school?
ygjb 9 hours ago [-]
Because that doesn't satisfy the narrative of American exceptionalism. It's easier to point at something and say it was stolen or copied than it is to compete, especially with the political climate in the US.
This isn't an anti-American sentiment. It is an anti-corporate/regulatory capture/embrace and extinguish sentiment (which probably reads the same to many people these days).
Gigachad 14 hours ago [-]
If the legal system declares the first thief’s theft not theft then all bets are off.
BeetleB 14 hours ago [-]
> If the legal system declares the first thief’s theft not theft
But they didn't find it. The Big LLM provider accepted guilt and paid a fine.
You can argue whether it was a fair amount they paid, but there is no legal precedent that was set. It's still considered theft.
kennywinker 14 hours ago [-]
As i understand it, they accepted guilt for downloading stuff illegally. They didn’t accept guilt for incorporating all of human output into their model without consent.
TheOtherHobbes 12 hours ago [-]
Copyright law only considers illegal ownership of a work, so the crime - or tort - was making/acquiring copies without permission or payment.
Training from copies has been ruled fair use because it's "transformative" and not simply "derivative."
This is obviously debatable, but that's where the debate is at the moment.
michelsedgh 12 hours ago [-]
So basically because they just browsed and used the information that was mostly public on the internet and they didnt copy it, they just learned from it and thats fine. Which makes sense. None of the llms let u copy someones work exactly anyways... makes total sense honestly. So in this case what happens to distilling? Is that also learning or ur trying to get to their actual weights by kind of reverse engineering it? Where would the argument fall there?
kennywinker 8 hours ago [-]
> but that's where the debate is at the moment.
Because of the rulings of a couple of judges. Is that actually what the majority of people think?
> Copyright law only considers illegal ownership of a work
That's definitely not true. File sharing, for example, is illegal even if you legally own the original copy you're sharing.
Similarly, copyright has something to say if I read a legal copy of harry potter and then create a new work in that world.
BeetleB 10 hours ago [-]
> They didn’t accept guilt for incorporating all of human output into their model without consent.
Because that use case is actually permitted by law.
kennywinker 8 hours ago [-]
I mean... that's one interpretation of the law, sure.
The law was written before the idea of an LLM existed, and some judges in some specific cases decided the previous law covered this usage.
So, it comes down to if you believe a couple judges ruling on a couple cases is the right way to determine a world-altering new legal framework.
nolok 14 hours ago [-]
> But they didn't find it. The Big LLM provider accepted guilt and paid a fine.
That's not how it works. You have to give it back.
Otherwise, the distiller can just pay a fine (no larger than the original did) and be okay then, right ?
BeetleB 10 hours ago [-]
[dead]
throwaway27448 13 hours ago [-]
True. It's the courts that failed humanity. Or perhaps the shits that invented copyright to start
mannanj 14 hours ago [-]
Is it theft if another thief steal's the first thief's theft?
itemize123 5 hours ago [-]
question's phrasing made your judgement obvious
miki123211 12 hours ago [-]
I don't think there's a ceiling to LLM capability. I do think that many software dev tasks are just far below that ceiling, and the gains to most dev work won't be that large from now on.
Where the new generation of LLMs (Fable, Sol) shines is tasks that are much harder than typical soft eng, yet that still have a verifiable answer, think mathematical proofs or exploits. I think there's still a good amount of low-hanging fruit in those (and similar) areas.
The next frontier after that is tasks that don't have automatically-verifiable answers, and may not even have correct and incorrect ones in the strictest sense of the word.
Reasonable lawyers might disagree on the question of "which trial strategy do I use given the following set of facts." There are answers that are clearly wrong, but being able to choose between many plausibly-correct ones requires many years of lawyering and seeing many trials play out. I do suspect that most lawyers are far below the ceiling that a hypothetical immortal lawyer that has practiced for an infinite amount of time would have achieved.
CMay 13 hours ago [-]
Or are humans more of a bottleneck than before, because to improve on the most complex problems that demonstrate intelligence you need some way to verify that they are correct. If it's hard for humans to even know if something is correct, wouldn't that slow everything down and simply put limits on the scaling speed of models based on human verification?
So instead of relying heavily on human bottlenecks, you focus on agentic task verification since that's the low hanging fruit and verifiable at scale?
edg5000 9 hours ago [-]
Very interesting point you make! Before LLMs I had a theory that we cannot make something more intelligent/complex than us.
LLMs are certainly more knowledgable, but maybe not more intelligent, arguably. It's possible we're approacing a ceiling indeed.
Model capability might be on an asymptote appraching but never quite reaching parity with human intelligence.
sscaryterry 14 hours ago [-]
Don't say that too loud, you may burst the bubble prematurely.
rllearneratwork 14 hours ago [-]
the ceiling is to eval's quality
Zambyte 15 hours ago [-]
Qwen 3.6 27b is already a viable default. I'm running it on a single 7900 XTX right now for Go development with pi. It's great.
bitexploder 14 hours ago [-]
I find 35B A3B viable as well, but your harness and runtime really matters to get tool calling and such dialed in. In fact, I would encourage you to experiment with it some as I find I get more reliable output from 35B A3B, though 27B is still generally smarter. A3B with a review cycle or two from 27B is great for me.
One of the reasons is, with good specs and design, A3B is just so fast. It isn't as smart as the 27B model, but it is close enough it can usually figure it out with the right tools.
monster_truck 14 hours ago [-]
Same! The only reason I'm not using it more is because it's summertime. I'm not in any hurry.
Setting the memory to "fast timings" is good for 8-12% more tokens/second if you haven't tried yet. I miss the slightly older days of AMD when powerplay tables were unlocked and we could configure the timings and voltages manually, there's another 30% being left on the table ez
MrDrMcCoy 9 hours ago [-]
What do you mean by 'Setting the memory to "fast timings"'? The only runtime I can get working for my GPUs is llama.cpp, which I haven't seen anything like that in its argument set. My perusal of the options for vllm and sglang didn't suggest anything similar either before failing miserably.
snapplebobapple 15 hours ago [-]
Works great with room to spare on my lenovo pgx too
icedrift 15 hours ago [-]
I'm still skeptical of the smaller models after the talent exodus a few months ago.
jimbo808 15 hours ago [-]
At this point I feel like the only factor differentiating SOTA models now is who they’re propagandizing you on behalf of (not considering agentic tooling/state management, etc).
dw_arthur 11 hours ago [-]
Who is going to break it to the Americans that China is more than a slight favorite to win an existential battle over which country is better at math?
d2p 16 hours ago [-]
I clicked through and it showed Qwen at the top at 55.4 compared to 55.3 for Opus Max. I have a screenshot.
Then I clicked away and back, and now it goes Qwen second, with 58.4, to Opus Max at top with 59.2.
I have screenshots of both. The description above the chart is the same in boh cases:
> Artificial Analysis Agentic Index
> Represents the weighted average of agentic capabilities benchmarks in the Artificial Analysis Intelligence Index (GDPval-AA v2, ³-Banking)
What happened? How can the scores change so much in a few seconds?
> HLE, AA-LCR and AA-Omniscience are now graded by GPT-5.6 Luna (medium), replacing GPT-4o, Qwen3 235B A22B 2507, and Gemini 3 Flash Preview respectively. These checks are now unified under a more capable modern model, selected for strong agreement with human judgment in our grader validation
Interesting that they chose a nano-sized model from OpenAI to be a grader for benchmarks involving knowledge and hallucination.
nolok 14 hours ago [-]
What's interesting is that if you ask 5.6 Sol or Opus 5 they will tell you it's a bad idea to have the reviewer be the dumber of the set as it can't judge them properly to decide who is right, and thus if one is better because it found an answer that's better but contradict the obvious it would be biased against. I know because I just had a consensus conversation with them this afternoon about a design that was similar (though about something completly different than judging agentic quality or whatever).
ahartmetz 16 hours ago [-]
Fixed the result, eh? In both senses of the word.
splatzone 15 hours ago [-]
Can someone please explain what changed, when it happened, and whether it was surreptitious?
johnnyApplePRNG 15 hours ago [-]
I have been suspicious of these AI leaderboard sites for some time now, and this only increases that suspicion.
torginus 15 hours ago [-]
In that case they should clearly label that this is a new benchmark.
gpt5 15 hours ago [-]
What was the change?
Gcam 13 hours ago [-]
Hey! George from the Artificial Analysis team here. We published an update today that does result in a change of the order, Qwen3.8 Max to second rather than first. The methodology change was an already planned upgrade to our equality checking/grader models, and brings the latest ³-Banking version to Artificial Analysis. Regular updates are normal for us to keep our benchmarks up to date.
The order changes but I think the story discussed in this thread holds - this is a very impressive release and Qwen3.8 Max is a huge step up in agentic capabilities.
> You gotta admit the timing looks very suspicious.
Do you mean the timing looks like: "We're SV tech-bros. Our benchmarks showed a chinese model above what's considered the best model at the moment. So we quickly modified the benchmark so that our SV tech-bros don't look like they're losing to a chinese model"?
That's indeed a bit fishy.
personjerry 15 hours ago [-]
They should probably freeze the results before publishing.
apitman 16 hours ago [-]
Welp. That didn't last long
WD-42 16 hours ago [-]
Same, they just updated it. Hacker news effect?
eli 17 hours ago [-]
I believe it. It's extremely good at troubleshooting. I gave Qwen and Kimi K3 the same annoying, complicated, intermittent bug to track down. Kimi did a bit better in understanding the existing code, but Qwen built some diagnostic tools and did an excellent statistical analysis on the log data. Qwen got way closer to the truth.
I'm very much looking forward to their forthcoming smaller model Qwen 3.8 releases. A version that can easily run locally would be great.
thefourthchime 17 hours ago [-]
Did you also try Opus 5 and 5.6 Sol?
ghosty141 14 hours ago [-]
5.6 sol was very impressive for me. I had a weird behavior while using Qt and I gave it a screenshot and my expectation of what should happen and it read the Qt sourcode and showed me that my issue was a bug (including link to the ticket).
sscaryterry 14 hours ago [-]
Opus 5 is just terrible
comboy 17 hours ago [-]
How CLI are you guys using for qwen and kimi?
eli 17 hours ago [-]
I use https://pi.dev/ which works fine out of the box but is fairly minimal and intended to be customized. There are many extensions.
OpenCode or oh-my-pi might make more sense if you just want a batteries-included agent. You can also make Claude Code work with other models without too much work, but I think that's asking for headaches.
trey-jones 17 hours ago [-]
I used claude with GLM and it's easy to set up, just hard to find the documentation. No headaches really, unless you want to use it against multiple different APIs.
Gooblebrai 16 hours ago [-]
Is there any subscription of any kind for Qwen? Or via Pi.dev needs to be used with API credits?
lkt 15 hours ago [-]
Opencode Go has Qwen 3.8 Max at $10/month
Gooblebrai 14 hours ago [-]
Found the usage limits on OpenCode Go quite poor tbh
Especially the second one seems exactly like my experience.
hungryhobbit 16 hours ago [-]
The cursing thing blows my mind. "User is upset? Let's make decisions even faster (ie. more wrong) because clearly that's what they want!"
It's a simple switch to make: cursing = try harder instead of cursing = stop trying. Is it really impossible to train Claude that way?
kloop 11 hours ago [-]
It's training data might have a ton of examples of people hurrying and screwing more after being yelled at
moffkalast 16 hours ago [-]
I'd certainly rank it at the very top of the want to kill yourself when using it benchmark. It outperforms everything else on that leaderboard.
With weaker models you can sort of understand, they're trying their best and failing, but this thing just channels its immense inteligence into being as annoying as possible instead. I know it can do what I'm asking it to do, but it just finds a way to weasel out of it, or maybe just thinks for 10 minutes instead, then fixes one thing and breaks four additional ones.
msp26 15 hours ago [-]
yep matches my experience completely
But even fable has the annoying tendency to invent new jargon and produce an incomprehensible soup of text.
cromka 15 hours ago [-]
Is there any model that knows how to smooth an overly literary text over? I find Opus and Fable constantly decorate the documentation they write like a damn 19/20th century writer. We're working with IT stuff yet it writes like it's going to win some Pulitzer prize. It's that one thing I don't get why they can't train them to do properly: I have not encountered a model yet that sticks to the current language of the domain it's tasked with.
msp26 15 hours ago [-]
Not sure how to fully fix this but I remember a session last week where I got so fed up mid way though reading a response that I used the following:
"give me this again without jargon invented this session at high density
and with a couple (maybe more or less) simple useful ascii diagrams underneath each design"
The context is that I was discussing an experimental new idea for my video game review analysis product.
Designs 1,2, and 3 were horrible: the model even suggested a rejection after the word soup so it would have been pointless to waste my fleeting time on Earth reading it.
Otherwise, I generally really enjoyed using fable for bouncing ideas. It was an absolute joy to have this thing provide useful criticism, analyse sample data, and create prototypes so that I could elevate my understanding of the problem without stepping down from a pure intuition/design headspace.
But I don't consider the purely model written code usable for a feature this important. I'll probably scrap it entirely and start from scratch with newfound understanding.
sscaryterry 14 hours ago [-]
I've ditched Anthropic completely because of it. It makes me furious.
copperx 17 hours ago [-]
I'm dumbfounded to see Opus 5 making SO MANY mistakes in coding simple stuff. Most times, Fable 5 comes out to be cheaper because it nails so many things much quicker than Opus 5.
garciasn 17 hours ago [-]
I have Fable plan and Opus implement. I haven't had any major issues working this way; however, Opus does seem plain fucking stupid compared to what I experienced with Sonnet previously.
aenis 16 hours ago [-]
I do the same, and generally have good results, but it does stupid things with gusto.
I'd open a blog with "weird things Opus did". Today it launched a swarm of cpu-hogging processes to test if the widget showing machine and I/O load is rendering nicely and correctly. The test went fine, but it was no longer able to kill those processes since they were really effectively hogging the CPU in various ways - being diligent, some of them were hogging CPU, some were murdering the SSD, some were pounding on the network adapters. Took me 30 mins to recover the machine to a working state without killing the meaningful, messy, in-flight sessions i had going on on other projects.
petesergeant 16 hours ago [-]
> however, Opus does seem plain fucking stupid
Infuriatingly so, in a way I don't remember Opus 4.8 being, but maybe I've just been ruined by Fable 5.
hbn 16 hours ago [-]
I bought my first LLM subscription with Claude right before they gave access to Fable 5.
I got so used to it, when they finally pulled access for me and I had to go back to Opus I felt like I was working with my hands tied.
I finally know what those women with AI boyfriends felt like when their app updated and it won't dirty talk with them anymore.
moffkalast 16 hours ago [-]
Fable has spoiled us all.
sscaryterry 14 hours ago [-]
Not so sure, I'm sure Opus 5 is just shit.
moffkalast 5 hours ago [-]
Eh it's better than 4.8 in terms of what it can get done on a good day, it's just far more taxing to get it there.
Like the Fable ban stunt, I wouldn't put it pass Anthropic to kneecap Opus deliberately to drive more people to their more expensive option.
usef- 16 hours ago [-]
Weird how different people's experiences are. If it's making simple mistakes something must be wrong in your setup/context I assume? It's been solid for me, beyond the usual LLMisms that all models have. But I keep context pretty minimal.
efficax 15 hours ago [-]
Every model that comes out comes with a bunch of people saying "this one is actually dumb they were smart before" and I don't really get it. The models since Opus 4.5 have all been basically the same to me. Sometimes they do the wrong thing, so you have to steer and stop and correct them. Leaving them to operate on their own in no-human-in-the-loop harnesses often gets bad results. But if you single thread it, and keep your work targeted (you have to know what you want the thing to do!), clear your context, the models will do what you ask pretty reliably.
PacificSpecific 11 hours ago [-]
Glad to see this comment as this has generally been my experience as well. I'm really curious to see why it's so infuriating for others. My best guess is I'm using it more conservatively than most other users in this thread.
cromka 16 hours ago [-]
Statements like this typically come from working on the same setup and context using different models. I actually have that very experience now; I work on something security-adjacent so Fable often drops out, at which point Opus behaves like its lobotomized half-sibling. Pardon me the language, but I can't find a better example to be honest.
yeeeloit 11 hours ago [-]
At this stage in the game almost none of the comments or articles on HN can be trusted, if you know what I mean...
nimonian 16 hours ago [-]
Agreed. Opus 5 is doing just fine, slightly better than 4.8. It's personality is insufferable, but I find myself catching fewer problems at code review. It generally understands my conventions and isn't so eager to accrue tech debt.
TacticalCoder 15 hours ago [-]
> I'm dumbfounded to see Opus 5 making SO MANY mistakes in coding simple stuff.
To me it's not so much the dumb mistakes (although there are some of those) but the ultra-verbose, mega-inefficient "solutions" to some problems / prompts.
Stuff that "works" if you're the kind of person that considers slamming a semi-trailer at 200 mph into a door did, technically, result in the door being somehow "open".
As it's supposed to be one of the most advanced model, I can't help but wonder if the solutions are that bad/verbose/inefficient because we're already in a loop of models being trained on sloppy-pasta from previous models.
visarga 17 hours ago [-]
Sent to solve one task, came back with half of it solved and 2 more problems.
capnjazz 16 hours ago [-]
"One thing worth your attention", "Two things worth knowing", "One thing to eyeball"
FridgeSeal 15 hours ago [-]
And one of them is always something just completely out of scope and the other is something obvious it missed.
“One thing worth your attention, if you were to detonate a pipe bomb in your house, it would have a negative effect on your living room”.
ethin 13 hours ago [-]
I don't use Claude code, just Claude web, and I get this all the time. Or (since I have it push me to actually think) it will ask me some question in our back-and-forth, and then right after it'll provide the answer. As a "hint". Like come on
greenchair 16 hours ago [-]
This is driving me crazy. opus 4.8 did not do this to me not (at least during pre-5.0 timeframe). Feels like the new cycle is one step forward, two steps back.
vunderba 14 hours ago [-]
What really enrages me is the amount of effort it puts into justifying weaseling out of work. (THAT'S MY JOB!)
It will do everything it can to defer or push it off, to the point where I’ve had to add multiple imperative directives to the AGENTS file telling it, in no uncertain terms, not to defer tasks under any circumstances.
cyanydeez 13 hours ago [-]
sounds like someone needs a local llm.
vunderba 13 hours ago [-]
Oh I do. Headless 128GB RAM machine serving llama.cpp with a number of local models that I use on a daily basis.
• Qwen3-VL picks up new images in a NAS, auto captions and adds the text descriptions as a hidden EXIF layer into the image, which is used for fast search and organization in conjunction with a Qdrant vector database.
• Gemma3:27b is used for personal translation work (mostly English and Chinese).
• Some small 8b models (like llama3.1) for sentiment analysis on text.
But haven't really tried using local LLMs in conjunction with agentic harnesses yet.
cyanydeez 12 hours ago [-]
recommend opencode w/qwen 35B or 27B with MTP.
My secret sauce is to use LLAMAcpp's reasoning-budget and reasoning-message that trigger cut off to overthinking with a message that says to either us subagents or compress the context. opencode's dynamic context pruning plugin can get you pretty far into the stratosphere.
> The Sleev (the project has been renamed to make a startup) creator was shilling their project in the OpenCode Discord. That person is very convinced they have something that no one has ever built before. They focused on token reduction without any real evals for capability impacts.
I'm generally against this context pruning without prompting or details. Sleev is very opaque about how it works and definitely will bust your cache.
vunderba 12 hours ago [-]
> My secret sauce is to use LLAMAcpp's reasoning-budget and reasoning-message that trigger cut off to overthinking with a message that says to either us subagents or compress the context
Thanks for the tip - I like this a lot. I remember having to do a lot of tweaking to curtail Qwen QwQ-32b when it would go down an endless psychotic recursive reasoning loops as part of its "chain of reasoning."
cyanydeez 1 hours ago [-]
I tailored the agent to ao both its system prompt and budget-message align.
I havent yet tailored the pruning messages, but mostly it works.
Reasoning budget can also be set by client, so potentially smarter.
Fordec 16 hours ago [-]
Yeah, I've dropped back to 4.8 entirely for the remainder of this billing cycle. I'm going to be seriously looking into Qwen adoption and harness migration options over the course of August.
thomasfromcdnjs 15 hours ago [-]
Same.
I could not get Opus 5 to do anything without losing a few years of my life from stress.
Fable has been okay but I am doing ML work and not allowed to use it which feels insane.
CuriouslyC 15 hours ago [-]
Ironically, Opus 5 is the most benchmaxxed model I've seen from Anthropic. It is legitimately smart in a lot of ways but it has communication issues, both in terms of how it communicates (all the autism of GPT class models, without the brevity) and how well it catches all the nuance of what you tell it.
I don’t see anyone talking about how you have to completely change your prompting strategies with Op. 5 versus 4.8 to get the most success.
combyn8tor 14 hours ago [-]
It works fine for me. Only issue I have is that it has me constantly reaching for the dictionary.
enraged_camel 16 hours ago [-]
It's my daily driver. I like it and find it noticeably better than Opus 4.8.
After I started reading complaints about Opus 5, I gave Fable the task of evaluating a bunch of code Opus 4.8 had written and compare it to Opus 5's code. Fable ran a dynamic workflow and the scores came back 15-20% higher for Opus 5's code in terms of quality, correctness and readability/conciseness. I did not tell Fable which Opus wrote which code, and I turned off memory as well to ensure there was no pollution from that angle.
My only complaint is that Opus 5's prose is annoying as hell. I wrote a custom skill for it for concise debriefs and it has been working pretty well for me.
agopaul 6 hours ago [-]
> My only complaint is that Opus 5's prose is annoying as hell. I wrote a custom skill for it for concise debriefs and it has been working pretty well for me.
I’m doing the same right now, and I’ve found that asking for “simple English” works most of the times, although not always.
Did you find better wording that works consistently?
cesarvarela 15 hours ago [-]
It is infuriating to interact with, but it is also first in many blind test leaderboards on LLMArena
fellowniusmonk 16 hours ago [-]
I have some internal tests I use for areas where one particular solution/paradigm is dominant but worse.
Opus 4.6 is the last model that's actually useful and can "adjust" its perspective to use the newer & better solution.
Where Opus 4.8-5 has over fit training on worse/older but "dominant" solutions it refuses to adjust.
Not only does this create an existential threat to adopting progress but it also means that if you have a code base that has rare but real world tradeoff the newest versions of Opus 4.7, 4.8 and 5 are worse than useless and become a major dev timesink.
nomel 16 hours ago [-]
What's the clear best, that you see?
dgellow 14 hours ago [-]
Hilarious to see only different responses
ofjcihen 14 hours ago [-]
Should have been “clear best and what do you do”
drschwabe 16 hours ago [-]
GPT 5.6 Sol
kachnuv_ocasek 16 hours ago [-]
GLM 5.2
petesergeant 16 hours ago [-]
Fable 5
logicchains 17 hours ago [-]
"As you requested, I've finished task X. Honestly, task X turned out to require task Y, which I haven't actually done. Task Y is the next step if you'd like to continue along this route."
pornel 16 hours ago [-]
This is the hard-won load-bearing quote.
dr_dshiv 15 hours ago [-]
Belt and braces all the way down
paradox460 10 hours ago [-]
The shape of this problem is very heavy
vardalab 15 hours ago [-]
Yeah, I told it to save in its memory that I don't want to have any more word
salad!
sunaookami 15 hours ago [-]
Can not confirm, for me it's the complete opposite.
This jumps around a lot based on the top throughput and latency of whatever provider happens to be best at the moment.
d4rkp4ttern 16 hours ago [-]
All these "intelligence" benchmarks miss something extremely important when using an LLM in a code-agent harness: How it communicates with you about what it did.
Opus-5 is practically unusable (for complex tasks) in this sense - its updates are voluminous, and dense with cryptic language (there are numerous reddit threads complaining about this, so it's not just me). I often have to ask it to re-state concisely in plain terms.
For a fairly gnarly task, after fighting with with Claude-Code + Opus-5, I ported my session to Codex + GPT-5.6-sol, and it was like a breath of fresh air.
Arguably a key aspect of intelligence is concise, clear communication, and current benchmarks miss that, at least as far as I'm aware. I would think some arena-type benchmarks where humans rate responses would measure this, though I'm not sure which those are.
lol. I am being downvoted for trying to help people out.
This community is pure trash.
HDBaseT 13 hours ago [-]
Instead, you should say something like:
"You can adjust the output style in your '.claude/settings.local.json' file".
OR
"You can decrease verbosity by doing x, y and z."
chpatrick 15 hours ago [-]
Is that what it feels like when the models get smarter than us?
msp26 15 hours ago [-]
No the models are just ass at communication without being directed.
Try asking them to make useful diagrams for some stuff in a codebase, out of the box without excessive hand holding they don't make good choices about what's worth communicating and how to do it.
You see this in their pointless frontend copy all the time too.
embedding-shape 14 hours ago [-]
Like any time you make them do any UI without strict directions they'll almost always add a label describing the feature somewhere. Ask for a calculator, and instructions for what the different buttons do might appear in the bottom out of nowhere for example.
Same concept of "over-sharing" seems to prevalent in a bunch of domains when it comes to LLMs, sometimes more visible, sometimes less.
cindyllm 13 hours ago [-]
[dead]
gpt5 15 hours ago [-]
A smarter model would know how to communicate with you correctly, and not just throw jargon it has just invented at you without explaining it.
chpatrick 15 hours ago [-]
But if you have two experts in a field talking to each other you wouldn't expect them to dumb down their communication.
cloverich 15 hours ago [-]
Concise, jargon free or limited explanation is the opposite of dumbed down. It requires to most skill and understanding to do well. Opus 4.8/5, for whatever reason, are getting worse at this crucial skill.
riknos314 15 hours ago [-]
Effective jargon usage is understood by the target audience.
If the AI is communicating to me and can't select the appropriate jargon level, it's failing at communicating effectively.
chpatrick 14 hours ago [-]
Or you're below its level.
ranguna 5 hours ago [-]
That's a loadbearing, heavy shaped, second take worthy statement
fuck_google 13 hours ago [-]
[dead]
IanCal 15 hours ago [-]
s/model/engineer
computably 15 hours ago [-]
"There is a view in some philosophical circles that anything that can be understood by people who have not studied philosophy is not profound enough to be worth saying. To the contrary, I suspect that whatever cannot be said clearly is probably not being thought clearly either."
micw 15 hours ago [-]
Guess that's the exact point of the "intelligence" benchmarks
HappyPanacea 15 hours ago [-]
No, a smart model should also give a concise executive summary, "brevity is the soul of wit".
a2ff6eeb0 15 hours ago [-]
So, in short, once models get smart enough they stop bothering telling us what they did.
Yeah, makes sense. A parent wouldn't bother explaining the details of their job to a toddler.
FridgeSeal 15 hours ago [-]
Which would be fine, but the parents also just smeared tomato sauce over the walls too, so let’s not get too ahead of ourselves.
jiggawatts 15 hours ago [-]
GPT 5.6 has similar language quirks that makes its comments nearly unusable.
I wonder if this is a side effect of MoE models — they can write excellent prose, but not simultaneously with writing code.
notfromhere 13 hours ago [-]
I think it’s just where they focused RLHF resources. The models have generally only gotten worse at writing.
And writing doesn’t have validators like code so you can’t really scale it in the same way
octoberfranklin 14 hours ago [-]
> and dense with cryptic language
Yeah if you ask it a medical question, it answers in that impenetrable jargony style that clinical journals use... full of unnecessarily custom adjectives ("orthopedic" instead of "of the bone") and discipline-specific terms (anterior, distal) even when the user didn't display mastry of this terminology (hint to frontier labs: add training cases for this; it will improve your model's usability).
My theory is that LLMs perceive the writing styles of various fields as being like different (but related) languages, and they're inclined to answer a question in the language of its source material unless specifically asked otherwise. If you add "ELI5" the model treats it as a question plus a translation task.
I think this is why programming questions are answered with an exaggerated cringey form of HN-speak ("load bearing", "gate" as a verb, "dissolves") by some models.
plaguuuuuu 6 hours ago [-]
it's like asking a developer to explain something.
I always get Haiku to rephrase anything human-facing.
moffkalast 16 hours ago [-]
Damn I thought it was my extra instructions, I swear everything it writes is in some shorthand with direct references to variables that literally nobody could figure out unless you literally just wrote that code 5 minutes ago. I had it stop writing comments altogether cause it was always four lines of complete and utter nonsense, and it doesn't even obey that rule half the time. Despite doing an extensive back and forth to make a complete plan, 5 seconds into the implementation it changes its mind and makes another assumption, adding some extra thing that tends to break the entire approach and needs follow-ups to repair or cleanup. Instruction following is basically non-existent compared to Fable, it just does whatever the fuck it wants.
fearmerchant 16 hours ago [-]
Everything is load-bearing with 3 measured blockers.
pixelready 15 hours ago [-]
Don’t forget the smoking guns! I think these new models have been reading too many Agatha Christie novels.
satvikpendem 15 hours ago [-]
Eh I don't know, I care whether it gets the job done and I can see the difference when I review the code, not how well it needs to explain the code to me, I can just read it myself.
hungryhobbit 15 hours ago [-]
It seems like if latency is having such a big effect that it's changing the winners, maybe your tests are awful and shouldn't be so latency dependent?
I mean, I get it: how fast a model responds is relevant. But a test that changes second by second is far less relevant than a test that tells you how smart the model is, and accounts for latency in some way that isn't constantly changing the result.
seizethecheese 14 hours ago [-]
Latency isn't changing the results for the coding index or arena ELO, but neither of those take latency or throughput into account, so we added those to our leaderboard as score components.
Latency and throughput matter a ton as a user, so I think it's actually totally defensible for a leaderboard to bounce around a lot as these numbers change. The best model to use changes a lot based on these!
They didn't run all benchmarks. It's the best in AA agentic index (GDPval-AA v2, ³-Banking) but not coding index (DeepSWE which is missing, Terminal-Bench v2.1 they have 81% vs 90% for Sol, SWE-Atlas-QnA missing).
moritzwarhier 17 hours ago [-]
Does "artificial analysis" mean what it says? Dubious.
But: I've been very impressed by the larger Qwen Models, and a brief try of Kimi also impressed me.
A lingering sense of quality degradation when going deep remains.
But that's not an accusation: they seem to be hitting the compute/quality tradeoff extremely well.
And on-prem capability is simply irreplaceable.
Apart from all the innovations that were driven by the strive for this optimization: quantization, "distilling" (without obvious mad-cows-disease)... I think China was an invaluable player in this progress. Intuitively, I'd even go so far to speculate that LLaMa wouldn't exist without the competition.
amelius 17 hours ago [-]
According to those graphs, Grok 4.5 appears to be the most cost-effective model.
user43928 16 hours ago [-]
$0.05 per task, Intelligence Index score 52 -> GPT 5.6 Luna max
$0.36 per task, Intelligence Index score 56 -> Grok 4.5 high
$1.13 per task, Intelligence Index score 58 -> Qwen 3.8 Max
$0.81 per task, Intelligence Index score 59 -> GPT 5.6 Sol xhigh
$1.80 per task, Intelligence Index score 63 -> Opus 5 xhigh
scrlk 17 hours ago [-]
Different benchmarks:
> Artificial Analysis Agentic Index: Represents the weighted average of agentic capabilities benchmarks in the Artificial Analysis Intelligence Index (GDPval-AA v2, Tau³-Banking)
> Artificial Analysis Coding Agent Index v1.3 incorporates 3 benchmarks: DeepSWE, Terminal-Bench v2, and SWE-Atlas-QnA
Qwen3.8 Max is 55.4 on the Agentic Index but hasn't been tested for the Coding Agent Index.
apitman 17 hours ago [-]
Looks like coding agent is model+harness. There are far fewer models represented on that page. I believe "agentic index" is still the metric to look at for coding performance. I could be wrong about that though.
Even then, this seems a much more marginal win than the headline suggested to me.
theropost 17 hours ago [-]
Anthropic is a bit nuts, I had $260 of credits on my max account for the extra usage the other night. It was expiring, so I figured I'll fire up an agentic swarm to deep dive and make some deep changes to some old cold bases.. literally 25 minutes or less, $260 burnt, it didn't get get into the implementation, just wrote a ton of useless plans for the most part. It really opened my eyes to what they expect to charge people.. wayyyy overpriced.
tarnith 16 hours ago [-]
Hint: The new models are really good at burning tokens.
I've had to use it a bit for work, and it's been remarkable watching the degradation in performance with the default suggested current models (Opus 5 as a prime example) vs the models that got them huge attention a year ago (Opus 4.6)
If you give 4.6 a spec, or existing code to implement a feature in, it will ask some pointed questions if there's something unclear in the spec, and then produce a plan and move to implement it.
5 will freak out at even a basic task, ask itself if it's own assumptions or your instructions are correct, proceed to re-assess it's own plan, and it's instructions 3-4 times, and then maybe produce code after burning several hundred thousand tokens (and quite a bit of time) analyzing existing code and thoroughly sweeping it for irrelevant problems both to the task it was given and the spec it came up with.
It's quite bizarre to me how well advertised the benchmarks and anecdotes from people one shotting MVP browser games are, compared to the experience of everyone I know that's had to actually use it to accomplish even a relatively basic task.
aenis 16 hours ago [-]
I managed to lose around $300 in credits I had saved for some emergency /fast sessions the following way: switch to Fable. Work on the design. Downgrade to Opus for the build. If any of other parallel Opus session has /fast enabled it seems to enable it for the newly spawned session by default. Before I knew it, the $300 was gone. I think the bug is now solved, but it was rather unpleasant. I dont ever remember bugs that would drain my wallet - with claude code its just another Tuesday. Still love it.
gnull 16 hours ago [-]
Claude code is just pool quality. They don't make how this thing will behave clear to the user, or give control. They fail at anything that needs an abstraction or model, not just APIs and shell scripts glued together. And "just ask AI" seems to be the default fix.
That vibe coding they brag about as if it was a good thing, it shows.
Take their notation for describing permissions. The docs are not comprehensive, and in practice it doesn't quite work how they describe it.
Or their management of sub-agents. I once lost a sub-agent, it finished and disappeared from UI. Apparently, you can't bring it back yourself: you have to ask the parent agent to do it for you. But the parent was Fable, and I ran out of credits, so I was locked out of using my opus sub-agent because of it.
Or an even more grotesque example: when you paste your claude API token to authorize, it covers characters with *. But it seems like an LLM has hallucinated a limit of API key length and the tail of your key stays visible.
hungryhobbit 16 hours ago [-]
What amazes me is how, for a vibe coded product where all they have to do is use their AI to fix things ... NOTHING EVER GETS FIXED!
I've probably gone to file 20 bugs. In all 20 cases there wasn't just one issue already filed for it: there were several, each which had a bunch of upvotes. And in all 20 cases ... every. last. one. ... Anthropic closed the ticket with no comment.
IF YOU ARE GOING TO HAVE A SHITTY VIBE CODED PRODUCT, AT LEAST USE YOUR SHITTY AI TO FIX THE SHITTY PROBLEMS!
cindyllm 15 hours ago [-]
[dead]
thejosh 16 hours ago [-]
so many ridiculous "how the fuck did this get through basic QA?" issues with Claude Code.
I can't believe how many critical bugs fall through.
My favourite one is the bug where Plan mode can execute destructive commands inadvertently.
Then all these get closed with `Closing for now — inactive for too long. Please open a new issue if this is still relevant.`. Awesome.
ethin 12 hours ago [-]
I mean. This is what happens with vibe-coded projects. When there's no actual software engineering going on, I wouldn't expect anything better than this.
formerly_proven 16 hours ago [-]
> I can't believe how many critical bugs fall through.
Almost like CC is 100% vibe coded.
tempest_ 16 hours ago [-]
I dont love it.
Opus 5 is just a token burner.
I use fable plan and spawn opus 4.8 workflows which seems to work alright.
ethin 12 hours ago [-]
For me it's the opposite: I don't have $200 plus to throw at Anthropic every Month, and when I do get to use Fable it rips through my usage credits like there's absolutely no tomorrow.
Of course, the hilarious thing to me is that Anthropic likes to claim that the usage limits are because of resource allocation problems or something like that. Obviously no such issue exists, otherwise they wouldn't allow you to bypass it by just paying a bit more and it would be a hard limit. So usage credits are entirely their way of just screwing you out of more money.
aenis 16 hours ago [-]
I suspect it must depend on how one manages their codebase - wrt to docs, ADRs, and general guardrails.
For me it is not great for design work - Fable is way better, and 4.8 was conservative and thus better (Opus 5 seems to jump to conclusions far more eagerly). But for overnight builds, where I give it 8hrs worth of work on LLDs created by Fable - its great. Where Opus 4.8 would often lose the plot and stop for questions clearly answered in the LLD - Opus 5 does manage to complete. Since it launched, I don't remember it ever disappointing me with builds. But designs? Boy, is this thing explosively stupid sometimes.
robbru 16 hours ago [-]
Opus 5 loves to stop working "for safety reasons" and shuts down the session! I avoid it at all costs now. Opus 4.8 has been my default as well.
16 hours ago [-]
mikae1 16 hours ago [-]
And at that cost they're still not profitable. It's going to be a bumpy road ahead...
arrowleaf 16 hours ago [-]
I thought they are making a profit on API pricing? A quick Google shows somewhere between 50-70% margins on API inference.
bakugo 16 hours ago [-]
API pricing is almost definitely profitable, but at this point I assume it's a small minority of their inference traffic compared to subscription usage, and unlikely to make up for the rest of their expenses on its own.
enedil 13 hours ago [-]
Why would you assume so when companies 150+ people can only use API pricing? My assumption is that more people use Claude at work than personally.
arikrahman 16 hours ago [-]
Meanwhile I can do all that and more with reasonix harness for Deepseek with a cache hit rate of 99%. And that's with unsubsidized American providers like cloudflare or Digital Ocean
tyre 16 hours ago [-]
People keep saying this but from what we’ve seen, Anthropic models are marginally profitable and earn back their costs over their lifetime. The company is burning money building the next versions and other ventures (e.g. verticals), but the models themselves have been profitable.
gamblor956 16 hours ago [-]
They're EBITDA profitable, not GAAP profitable.
swalsh 16 hours ago [-]
I think profitability is a matter of accounting. Inference is where money is made, but training is where money is spent. We keep getting new models every few months, but frankly the old models are still quite usable. I suspect labs will soon start specializing in expert models per use case so they can increase the lifespan of individual models, and change the profitability per model.
CuriouslyC 15 hours ago [-]
That's not the only reason to go to expert models. The more different domains you try to stuff in there, the more parameters the model needs to keep things coherent and not overload tokens in a way that induces errors. For example, if a model trained only on biology text sees "sonic hedgehog" there's no ambiguity, and this compounds for all the things that are "overloaded," in the training corpus, which turns out to be quite a bit.
an0malous 16 hours ago [-]
What’s the blast radius of this bubble popping? It’s all private investment still right?
bhewes 16 hours ago [-]
Two thirds of most of the DC builds are not compute. So it's a CRE play the last leg holding up that mess.
swalsh 16 hours ago [-]
Its tough to go from max account at home and pay per usage enterprise account at work with heavy usage limits... but the limits are there because pricing is insane. Feel like I'm in the $5 Uber rides phase at home.
hahahaa 16 hours ago [-]
The Chinese models are the public transport in the uber analogy. Once the price the goes up catch the bus!
swalsh 16 hours ago [-]
Lol perfect analogy. I'm still paying for claude because the quality is unmatched.
cortesoft 17 hours ago [-]
It’s crazy how different the credit cost and subscription cost are.
With the $200 subscription, I can have Fable on ultracode working for hours and not dent the usage limits.
notatoad 14 hours ago [-]
yeah, i tried out GLM-5.2 when the news was all full of hype for that, and it's fine... definitely better value that API rates for claude. but comparing the value i got from that to the value i get from a claude max subscription... claude is way cheaper.
AlexandrB 16 hours ago [-]
VCs are footing the bill for that $200 subscription.
riknos314 14 hours ago [-]
The $200 sub is customer acquisition cost to hook devs that then become the marketing team trying to get their company to bring in Claude (at the highly profitable API price).
ericd 16 hours ago [-]
They have something like 80% gross margins, are at a $100B/yr ARR, and are growing at 10x per year... If that keeps up, they're going to be doing more revenue than Google in a year ($400B ARR, 20% per year growth)
dexwiz 16 hours ago [-]
How can you sanely project the last 12 months forward? We have seen a huge uptick in usage. Last summer AI was a toy to most devs, now every enterprise developer I talked to uses it every day. Coding agent providers are surely going to hit market saturation in the near future.
ericd 14 hours ago [-]
Maybe, maybe not. Personally, I hope local AI eats their lunch so that the benefits are more decentralized and accrue more to society generally.
I don't think you're right about that last prediction, at all. And new use cases are opening up as these get smarter. I think things are going to get pretty weird.
But the point was that it really doesn't look like they're losing money on users, on average.
senordevnyc 2 hours ago [-]
Where did that $100B figure come from? I thought they were at ~10B at the end of 2025, so they're either not at 100B yet, or they're growing way faster than 10x / year.
16 hours ago [-]
ux266478 16 hours ago [-]
At last, a valid usecase for VCs.
dionian 15 hours ago [-]
i'll take it, just hope they dont rugpull us soon. im sure its coming
brynnbee 7 hours ago [-]
I had same experience with OpenAI. I have the $200/month plan and use 5.6 Sol all the time. What would normally use about 2% of my weekly allowance burned through $100 of credits in 40 minutes.
hahahaa 16 hours ago [-]
You plugged in a space heater on a roofless house.
There is some element of responsibility on the user to guide and monitor the model/harness and not let it rip to burn tokens.
criddell 16 hours ago [-]
> wayyy overpriced
Maybe they consider that hiring a person to do it would have cost at least as much and taken much more time, so paying them is a bargain.
echelon 16 hours ago [-]
Yeah, but now we can hire the Chinese instead for 1/100th the cost. It's an even better deal.
Plus we get to own, keep, run, do whatever with the model. We don't feel trapped. Moreover, it's something we can truly build on top of and own our own destiny.
Anthropic and OpenAI are the new Oracle (Oracle pre-AI; Oracle is even worse now). Expensive, feels like dealing with a lawyer, and not at all open. They just became infinitely less cool than they were a month ago.
The whole of our industry is going to migrate to open weights. We're smart enough to know this is the better deal and technical enough to be able to pull it off.
The only thing that might save these OpenAI and Anthropic in the near-term is an abundance of enterprise contracts negotiated with non-tech companies. They'll soak consulting firms and F500 companies for "AI" integrations.
criddell 16 hours ago [-]
> the new Oracle
I think that's exactly what they are going for - enterprise and government customers.
pvtmert 16 hours ago [-]
Anthropic is the new AWS.
Amazon's first principle is the Customer Obsession. Making customers happy.
Fun bit is that the human psychology rates personal looking fixes better than having no issues at all.
For example, AWS overcharges you, you contact support, and more or less hassle free they refund or issue credits. The customer feels appreciated, or at least got something "extra" or "special treatment".
Meanwhile, any other (small) cloud. Simple, no weird charges. Even _most_ of network egress is free. But, no reason to call support or feel "extraordinary". Comes out as "meh" against Amazon's "top tier" support model...
john01dav 16 hours ago [-]
Anthropic's constant changing of its mind leads to instability which leads to unhappy customers
riknos314 14 hours ago [-]
Aws is an infrastructure company that builds services on top of that infra to sell more of it at a higher margin.
Anthropic trains models on AWS's (and GCPs, and Microslop's) infrastructure, then skims margin off of selling inference also on the infrastructure owned by the other companies.
These are extremely different businesses.
axpy906 15 hours ago [-]
I’ve never gotten a refund from Athropic.
polishdude20 17 hours ago [-]
You should just spend those towards a cursor subscription.
petercooper 17 hours ago [-]
Hopefully this boils down to the smaller versions they've teased. In my experience, Qwen models are the closest to the "less knowledge, more intelligence" (yes, the two are hugely correlated!) ideal some tool-dependent tasks need. Even the 3.5 2B can be easily prompted to always lean on tools and not jump to false conclusions (although its actual coding skills are abysmal, as you'd expect).
quotemstr 17 hours ago [-]
> less knowledge, more intelligence
People produce such models by over-RL-ing smaller models on math and coding tasks. I've found the results capable of neither innovative work nor thinking outside the box. They're straight-A students raised by tiger moments who never let them play freely for hours in the dirt.
Perhaps you could say such models are skilled --- but intelligent? Not by my measure.
People and AIs alike need diversity of experience and a broad liberal arts education to see hidden connections between fields and make real advances.
DC-3 17 hours ago [-]
It's amusing to me that AI has become sophisticated enough that people have started being racist to it.
petercooper 16 hours ago [-]
I agree with you to an extent, but you have certainly given me food for thought.
Sticking to LLMs, they seemingly get their intelligence (whatever that really means) from building models rich with knowledge, so you could have a point. But Qwen models seem to be particularly good, even at small model sizes, at maintaining both their own knowledge while acquiescing to and integrating external information in the moment.
syntaxing 17 hours ago [-]
I am so excited for Qwen 3.8 27B. It’s a shame how slow prefill (~3-400) is on a strix halo but it’s such a good model for agentic tasks.
colingauvin 17 hours ago [-]
Prefill is survivable if you cache well. But what kills me is the context. Qwen 27 needs a ton of room for KV Cache. I guess not an issue on a 128 GB Halo or Spark, but if you are running of consumer/prosumer GPUs it's miserable to be compacting every 120k tokens.
tarr11 17 hours ago [-]
What type of agentic tasks are you using it for (eg how complex)?
syntaxing 16 hours ago [-]
For personal stuff, I use it with AnythingLLM. It replaced any Google search for me. For coding, I run opencode though I have been debating switching to Pi. I would argue it’s at Sonnet 3 level.
LoganDark 17 hours ago [-]
I find that 35B-A3B is much easier to run on my M4 Max (both prefill and generation)
markasoftware 17 hours ago [-]
It's well known 35b is much faster (on any hardware) and quite a bit dumber
dofm 15 hours ago [-]
This really very much depends on how you are using it, I think. If you intend to leave it to solve long context problems and write whole prototypes, the 27B is going to be much better.
But if you are sort of pair-programming with the model, the speed obviously matters and I think then the 35B is acceptably smart, and when it's wrong it'll be wrong much more quickly. It seems very good on SQL and PHP, and I assume on typical JS and Python.
I would rather work that way, so I hope they do produce a small MoE model.
LoganDark 1 hours ago [-]
I had no idea. Where can we learn stuff like this?
CamperBob2 17 hours ago [-]
How are you running it on a Strix Halo? The weights aren't out yet, are they?
13rac1 17 hours ago [-]
I interpret @syntaxing as meaning they are looking forward to running Qwen3.8-27B, but are frustrated by prefill times with other models, such as Qwen3.6-27B.
syntaxing 17 hours ago [-]
I meant Qwen3.6. Unsloth supposedly has early preview of the model and the VRAM requirement is the same so most people expect similar model size and type.
SwellJoe 17 hours ago [-]
I find that surprising.
I've been trying it on several projects and have found it's pretty sloppy. It leaves stuff broken, doesn't reliably write tests to check its own work unless explicitly prompted, misunderstands the assignment, etc.
It is smart and reasonably quick but not reliable.
superfrank 17 hours ago [-]
I've come to the same conclusion over and over with all of the Chinese models that have been claimed to be catching up with OpenAI's and Anthropic's frontier models (Deepseek 4, GLM 5.2, Kimi K3).
At their best, I think they're closing in on Opus and GPT, but they're incredibly inconsistent and the variance in output quality is much higher than the best from any of the Anthropic or OpenAI models from the last few generations. The only way I can describe it is that it feels like a lack of intuition with the models which means I find my self needing to write longer prompts or have more back and forth to get them to do what I want from them.
To give an example, I have a saved prompt that I use as a sanity check on some data I'm storing. It reads about 50 rows from a DB and matches them to the UI and makes sure the data is displaying correctly. I've been using this with GPT 5.5 and now 5.6 for a few months and running it a few times a week with no issue. Sometimes I'll run it multiple times in a single chat if I notice bad data (run it, fix thing, run again, fix another thing).
I recently tried to switch to using Deepseek v4 (first flash and then pro) and while both did the task just fine, both would do things like change the response format from one message to another in the same chat or randomly decide to omit things it didn't think were relevant. At one point I ran the prompt, fixed some bad data, and then said "Okay, I fixed row 7, run {prompt} again" and so it decided to leave row 7 out of the response. A few times the first message would contain a table and then the next run in the same chat would contain the data in a bulleted list.
None of those are major issues and all could be solved with a bit more rigor in my prompting, but for me it makes them harder to work with. Those examples are a bit trivial, I think they're the easiest way for me to illustrate the gaps I see with them.
dyauspitr 17 hours ago [-]
It’s because they’re doing some sort of combined score of intelligence, speed and cost. On pure intelligence it doesn’t even show up in the top 10.
quirino 17 hours ago [-]
A couple days ago they had published an overall score of 53 for this model, but that was removed and today it returned with a score of 56.
I wasn't able to find an explanation from them. Anyone knows what happened?
whwhyb 1 hours ago [-]
according to them:
> launch traffic hit our public API endpoint harder than expected, causing intermittent instability.
Why does an open weights model cost nearly the same as GPT5.6? $1.14 vs $1.23 on the cost index. Since you can't presumably run this on your own hardware given the model size and hence gain other things like privacy, I don't see any reason to move away from GPT at this rate.
eli 17 hours ago [-]
It's not enough that it's better?
Many providers will host it and will compete on price. It also can't easily be taken away because one company (or one government) decides they don't want it around any more. People can fine-tune it for particular workloads.
Art9681 17 hours ago [-]
They cherrypicked benchmarks. The ONE weighed benchmark where is beats Opus5 by 0.1 points is what was linked because that's how propaganda works. The Agentic Index that includes the full benchmark suite has it in 5th place.
Might as well use gpt-sol.
iAMkenough 15 hours ago [-]
The whole industry cherry picks benchmarks.
I stopped paying attention to self-published benchmarks when Apple started using those non-sensical performance graphs with "relative performance" as a vertical axis when announcing a new chip.
drnick1 17 hours ago [-]
> It's not enough that it's better?
It's barely better, and barely cheaper, not really enough to challenge the status quo IMO. Half the price for basically the same performance would be a much stronger value proposition.
Things change radically month to month. Nobody is remotely close to capturing the market or having any kind of stability over time. People move around quite a lot, often to sidegrade within a generation. Just playing fly on the wall with discourse would be enough to tell you all of this, even without the data to back it up.
SwellJoe 16 hours ago [-]
If anybody has, it's DeepSeek. But, with the promised price hikes, I'm sure that'll change. I'm guessing they're raising prices not because they're not making a profit at those prices, but because they're running into capacity problems and need to slow down until they've got more or risk providing poor service. For now DeepSeek Flash is the best deal going for API usage and its popularity makes sense.
Also, OpenRouter misses most of the usage of the US models, as most people are getting those from the vendor directly via subscriptions.
eli 16 hours ago [-]
That's got a significant selection bias. Claude and ChatGPT and Gemini and other subs do not go through openrouter.
ux266478 16 hours ago [-]
Not really, because that's not a unique aspect of any of those. It's true of all subscription services (that I'm aware of), as well as all of the free models. The selection bias primarily will be against models which be an outlier in the difference between openrouter users and total users, which is a much harder position to argue for any given company except for maybe Twitter.
You can argue there's a selection bias that openrouter users are less likely to display model loyalty, but it would still be a visible confounding factor if it was a statistically significant behavior. And it's not. Nor is there a visibly meaningful indication that people don't sidegrade between models. With every single data set, you're going to see that. You're also going to see it reflected in discourse, as I mentioned. Fact of the matter is there isn't a status quo in AI any more than there's a status quo in cars.
apitman 17 hours ago [-]
For one thing, providers of open models can't arbitrarily increase their prices without facing competition.
frereubu 17 hours ago [-]
But given the extremely low cost of switching, why wouldn't you use the cheaper one if they're comparable?
apitman 17 hours ago [-]
As low as it is, switching between providers on OpenRouter is still lower.
Qwen Max is their large model - over a trillion params. Similar to Kimi K3 in size. Qwen 3.8 27B is going to be more accessible to your own hardware. I'd say that Qwen Max is not approachable for the majority of people and companies to self-host.
17 hours ago [-]
criley2 17 hours ago [-]
GPT5.6Sol completes the suite in 70M tokens, while Qwen3.8Max needs like 145M tokens. So this is a case where models like Qwen 3.8 and Kimi K3 use a lot more output (reasoning) tokens, go a good bit slower, so they can ultimately achieve a better intelligence score than if they went more quickly.
There are a couple of frontiers (ok bad word, maybe categories) in open weight models.
These Qwen 3.8 and Kimi K3 style models aren't trying to win on price, they're trying to compete on intelligence and capability.
Models like Deepseek V4 Flash (updated this week) are $0.03 a task, or 50X cheaper than Qwen3.8/Kimi K3, and 100X cheaper than Fable, while offering stunning intelligence. That's a different frontier for competition, and perhaps one more interesting for someone who wants to see them compete on cost.
benjiro29 16 hours ago [-]
Why does an open weights model cost nearly the same as GPT5.6? $1.14 vs $1.23 on the cost index.
What cost the most in API. Input, Cached Input, or Output. There you have your answer.
Unfortunately, we have moved so much of the actual intelligence of models towards reasoning, what results in some models getting good scores, but this is because they are dumping a insane amount of reasoning tokens at the problem.
So a mid priced model, with heavy reasoning output, cost the same as a expensive model, with medium reasoning output.
Before the GPT Luna price drop of 80%, you actually had the same price if you used Luna High and Sol Low. With the difference that Sol Low was insane fast, and often way better code.
Do not look at the top score but more what is on the horizontal axis as you go down. Sol Medium is frankly, was the best performance for dollar, until that Luna price drop. I will even argue that despite the higher price, Sol Medium is still way better despite Luna Max being cheaper. Or Opus Low, one of the better values also.
What do you notice? Is that those models all have a high intelligence start point for their low setting. So that means they do not rely as much on output tokens aka thinking.
ecocentrik 17 hours ago [-]
Why should open weights correlate with cost? Cost correlates with the expense of running the model more than it does to the expense of developing the model.
Alpha3031 17 hours ago [-]
You said it yourself, model size and hardware. Big models cost more (good optimisation reduces things slightly, but they still need the hardware).
jazzyjackson 16 hours ago [-]
Running a large model on rented GPU is still meaningfully more private than handing your chat logs over to FAGA
TheCycoONE 16 hours ago [-]
The acronym is new to me: Facebook, Anthropic, Google, openAi?
efficax 17 hours ago [-]
it's a big honking trillion some parameters model. it's not cheap to run
gerdesj 12 hours ago [-]
My vague equivalent of the pelican riding a bicycle test (for a local model without internets) is to ask it: "Where is Yeovil"? I don't expect a totally accurate answer for obvious reasons but I do enjoy watching the accuracy improve.
Qwen3.6-27B-FP8 currently espouses (see below), which is not too bad. The directions are a bit mad but the mileage is about right and there is a helicopter manufacturer here and a RNAS (navy not airforce) museum nearby at Yeovilton. Cosford is in Shropshire which is not a million miles away.
I'm not sure what 盆地的 means but the river Yeo is correct ... OK ... "basin like" - again not bad, even if Chinese is not the first language here. The model understands that Yeovil is named after (or vice versa or at least is associated with) a river
Yeovil is the current form of Gifle (Saxon) which I thought meant "bend in a river" but WP is currently saying "fork in a river". My source is a local museum. There is a fork but was it there 2000 odd years ago? My hydrology skills say ... possibly
----------------------------------------------------
Q: where is yeovil:
Yeovil is a town in Somerset, in the South West of England.
It is located roughly:
25 miles (40 km) south-west of Exeter
60 miles (100 km) west of Bristol
140 miles (225 km) west-south-west of London
Yeovil is known for its historic market town center, RAF Museum Cosford (nearby), and as a significant industrial town, particularly during World War II for aircraft manufacturing (including the Wellington bomber). It sits in the盆地的 valley of the River Yeo.
zmmmmm 15 hours ago [-]
The fact that the Chinese models have caught up on benchmarks suggests to me that its likely we will start to transition now into much more of a brand war. It will be subjective qualities that drive our decisions more than measures of absolute intelligence. Already I am choosing models more because I like the personality or style of what they do than because I think they have the absolute highest chance of outputting the most technically correct answer to any given prompt. It will be very interesting to see how things evolve in this direction.
mindwok 14 hours ago [-]
For me now it’s simply cost and speed. With GPT5.6 and Fable (and respective open models since then) we passed a threshold where intelligence is sufficient. Now I just need speed of iteration and good prices.
imagetic 5 hours ago [-]
It's the first model I've used that makes me forget it isn't one of the big frontier players after the first prompt. So far I'm impressed.
aliljet 17 hours ago [-]
Is there a path to distill this model to do very specific things? Like a RAG strategy for a small (or even large) corpus?
Alpha3031 17 hours ago [-]
Depends on what you want to do. Some task specific models can be trained with a few ten or hundred thousand training examples so you can use a bigger model to produce synthetic training examples and then fine tune a smaller student model. I think that's the usual process. Whether you'd get acceptable performance this way depends, as mentioned, on what you're trying to do and what you'd consider acceptable.
teravor 17 hours ago [-]
once you are able to get the full probability distributions per token you can distill it on specific domains. distilling without that isn't generally a good idea unless you have invested millions in the requisite infrastructure.
bonoboTP 16 hours ago [-]
I distrust any benchmark where Opus 5 beats Fable 5.
h14h 16 hours ago [-]
This has me hopeful for Qwen3.8-27B!
ben8bit 17 hours ago [-]
Haven't tried this yet, but going to soon! I have to wonder what happened at Anthropic. We've cancelled our subscription in favor of OpenCode & Codex. Sol is just so good & OC goes so far for every $ spent. Claude's become a pain to work with - average output with an annoying personality. Who knew this would be an issue even a year ago? In any case, loving the stuff from the Chinese models!
tomComb 17 hours ago [-]
> an annoying personality
I was with you until there. Qwen and the OpenAI models are great, aggressive agents, but they’re not as good as the anthropic models for human interaction. They just don’t have the subtlety, understanding, or attention to detail.
colingauvin 15 hours ago [-]
Claude 4.5/4.6 - absolutely agree. Fable 5? From my (limited) testing, also reasonable to interact with.
Opus 4.7/4.8/5? Absolutely smug and antagonistic and preachy. I'm constantly fighting with it to stop fighting me and accept that I occasionally know better. It's really frustrating to spend so many tokens of such an expensive model arguing with it.
ben8bit 17 hours ago [-]
Really? I've heard so many other people complain about this recently. And maybe it's possible that it's the prompt style even. But interesting that it's not across the board.
ngl999 9 hours ago [-]
It's censored and it'll spread certain kind of narrative all over the world.
MrDrMcCoy 9 hours ago [-]
If you're using AI for narratives, you're using it wrong.
londons_explore 15 hours ago [-]
I just don't think you can combine speed, latency, price and intelligence into a single useful metric.
Clearly the weighting of those things depends on the usecase
camnora 16 hours ago [-]
Qwen is just crushing it overall. I regularly use 3.7-flash for everyday coding needs and it gets the job done.
brettgo1 16 hours ago [-]
Out of curiosity, what's currently the best model I can use locally?
daemonologist 16 hours ago [-]
With an unlimited budget, Kimi K3 (which is quite comparable to this Qwen Max imo). With a normal budget/a PC you might already have, probably Qwen 3.6 27B.
arjie 15 hours ago [-]
$500k - Kimi K3 (maybe $250k? Haven’t done this one)
$25k - DSv4 Flash
$4k - Qwen 3.6 35A3B Q5
$1k - Qwen 3.6 27B Q4
Some people prefer the sense over the MoE YMMV.
colingauvin 4 hours ago [-]
16 DGX Sparks can run K3 at a reasonable TPS. So that's $64k.
2 DGX Sparks can run DS4 at 1 million context with 50 TPS so that's $8k.
1 A4500 can run 35A3B. Those are about $1200 new.
27B actually takes more hardware to run than 35B because attention is done differently I believe and therefore KV Cache takes a lot of space. It will run on an A4500 but it's slow and context will be like 32k.
apitman 14 hours ago [-]
These numbers look about right based on my experiences as well. Though for a single user I think 2x DGX Spark (~$10k) runs DSv4 Flash fairly well right?
Fordec 16 hours ago [-]
Anthropic have a real fight on their hands now. The competition is no longer 6 months behind, it's 6 days. If this had come out two or three weeks earlier this would be an absolute market leader on both quality and timeline.
steve-atx-7600 17 hours ago [-]
curious about methodology. ive seen them post results for claude/codex when they only ran over benchmarks 3 times per model...
dangoodmanUT 15 hours ago [-]
I'm seeing opus 59.2, Qwen 58.4?
looksjjhg 17 hours ago [-]
That took what 2 years? I love how the chip ban made them more efficient
Footprint0521 13 hours ago [-]
Facts lol, now all the Chinese models are 1/40th of the cost for the same intelligence
brcmthrowaway 17 hours ago [-]
Could someone like Apple be playing the long game - Good Enough(tm) intelligence will eventually fit in our pocket and homes?
colingauvin 16 hours ago [-]
DS4 Flash Q2/Q4 mixed quant fits on a DGX Spark (a $4000 device which is not particularly unheard of expense for Apple customers), and is indistinguishable for me from Opus for my personal daily use/assistant benchmarks[0].
Indeed. I like using Macs mostly, and the bargain M1 Max MBP I am using for local LLMs is a fabulous experimentation platform and does loads of other stuff well, so I am in no rush, but if I reached the point of buying dedicated hardware for an LLM, I'd be looking at the DGX Spark machines.
LPisGood 17 hours ago [-]
Almost surely. Apple is extremely well positioned to take advantage of this over the next decade.
kyxsc 16 hours ago [-]
Apple is already doing this... they worked with Gemini to distill the model into a smaller one that fits on your phone. If you have iOS 27 Beta, you're already using this
notatoad 14 hours ago [-]
sort of. they have a local model, it does some things. they also have significant cloud infrastructure backing it, and most tasks are going to be sent off to the cloud for processing, not be handled by the on-device model. Siri is not on-device by any stretch of the imagination.
delduca 17 hours ago [-]
Go China!
sirbor 17 hours ago [-]
Qwen is the way to go
esafak 16 hours ago [-]
It is also the most expensive open source frontier model, per task; cf. Cost per Intelligence Index Task. If it is as good as the benchmarks indicate it bodes well for Qwen and China. For my part, I'll pass; it is not on the Pareto frontier.
atemerev 16 hours ago [-]
Well, that's the bad index then. It is barely usable in my opinion compared to other Chinese frontier models.
ramon156 15 hours ago [-]
which one of the other chinese frontier models is better?
dyauspitr 17 hours ago [-]
It doesn’t even show up in the raw intelligence index, so how could it possibly be the best?
proxyscore 16 hours ago [-]
Does it matter, it's all non deterministic bs ware and deepseek is eating the Americans lunch
What I'm really excited for is the 27B model. 3.6 is still the king of local, and if 3.8 makes the same improvements it could really legitimately make local viable as a default. I'd love to run a perpetual agent on 3.8 that's locally driven.
When I finally put $15 into Deepseek and it beat the brakes off Codex 5.5 on multiple rather complex projects without any of the obnoxious mistakes, I was sick to my stomach with buyers remorse. I couldn't believe I ever felt like I was getting my moneys worth at $200/mo. I wouldn't even use OAI's models if they were free and unlimited at this point, I'll happily pay for what I already know works. No reset bingo, no cache errors, no annoying shitposters as a primary source of info. Oh, and I still had $10 of tokens left
And yes, 3.6 is excellent locally. The rest of this year is gonna be awesome
I suspect that the models are genuinely close and that certain experiences get felt across providers but are inconsistent enough to convince people one is superior to the other. I for one have tried Deepseek on and off since my co-founder is fond of it and I've stopped trying now because I never have a good experience.
On reddit et al., people talk about LLM brands like their sports teams.
I think the first-party ecosystem moats they're all trying to build are exacerbating this tendency, as now people have a lot of learning time sunk in a company-specific option.
I switched to DeepSeek entirely once I decided to put 10 bucks on it and I realized that it could do whatever I was throwing at Claude or ChatGPT prior to that.
I recommended it to one of my friends, and he was surprised DeepSeek could solve task that Claude got stuck at. I was surprised at it too.
I know others that tried and were less impressed too.
There would be no reason to if you are in the privileged position where cost isn't an issue.
For the rest of us something that's 95% as good for 20% the price is a hell of a value proposition.
Self-hosting is the biggest reason.
I think that may be part of it.
LLMs can be autonomous to an extent. All of them need steering - which is why I feel they are more a superpower the more I am an expert on the subject matter.
The more you want it to be autonomous, than yeah, you may benefit from using the very best the industry has to offer, however slightly better it is.
But if you are always in the loop anyway, you may want to try DeepSeek. You will get similar results for a fraction of the price.
Don't know about that but your neighbors in Iran in early january happened to be "nice people" who just followed the orders to slaughter 30 000 unarmed civilians.
We could talk about the, what 600 000 deaths, including many civilians, in the Ukraine/Russia war.
Or we could talk about the number of nice palestinians killed since the beginning of the war in Gaza. Or we could go a bit further and talk about the joy and celebration in Gaza after their heroes brought back 200 hostages after having slaughtered 1200 civilians.
You may be living in a place that you think shields you from those but I know the ideologies behind these acts.
The fallacy of gray is just that: it's not true that there's always a nice middle ground and that there's no evil ideology out there.
Something something about the price of liberty being eternal vigilance. For there are people abusing your blind trust.
I think it's mainly due to poor education many receive and a very controlled media that suppresses information.
It's shocking considering how much money they spend on education compared to other nations.
Not sure what part of being charged guilty and paying a fine you see as "free".
This isn't an anti-American sentiment. It is an anti-corporate/regulatory capture/embrace and extinguish sentiment (which probably reads the same to many people these days).
But they didn't find it. The Big LLM provider accepted guilt and paid a fine.
You can argue whether it was a fair amount they paid, but there is no legal precedent that was set. It's still considered theft.
Training from copies has been ruled fair use because it's "transformative" and not simply "derivative."
This is obviously debatable, but that's where the debate is at the moment.
Because of the rulings of a couple of judges. Is that actually what the majority of people think?
> Copyright law only considers illegal ownership of a work
That's definitely not true. File sharing, for example, is illegal even if you legally own the original copy you're sharing.
Similarly, copyright has something to say if I read a legal copy of harry potter and then create a new work in that world.
Because that use case is actually permitted by law.
The law was written before the idea of an LLM existed, and some judges in some specific cases decided the previous law covered this usage.
So, it comes down to if you believe a couple judges ruling on a couple cases is the right way to determine a world-altering new legal framework.
That's not how it works. You have to give it back.
Otherwise, the distiller can just pay a fine (no larger than the original did) and be okay then, right ?
Where the new generation of LLMs (Fable, Sol) shines is tasks that are much harder than typical soft eng, yet that still have a verifiable answer, think mathematical proofs or exploits. I think there's still a good amount of low-hanging fruit in those (and similar) areas.
The next frontier after that is tasks that don't have automatically-verifiable answers, and may not even have correct and incorrect ones in the strictest sense of the word.
Reasonable lawyers might disagree on the question of "which trial strategy do I use given the following set of facts." There are answers that are clearly wrong, but being able to choose between many plausibly-correct ones requires many years of lawyering and seeing many trials play out. I do suspect that most lawyers are far below the ceiling that a hypothetical immortal lawyer that has practiced for an infinite amount of time would have achieved.
So instead of relying heavily on human bottlenecks, you focus on agentic task verification since that's the low hanging fruit and verifiable at scale?
LLMs are certainly more knowledgable, but maybe not more intelligent, arguably. It's possible we're approacing a ceiling indeed.
Model capability might be on an asymptote appraching but never quite reaching parity with human intelligence.
One of the reasons is, with good specs and design, A3B is just so fast. It isn't as smart as the 27B model, but it is close enough it can usually figure it out with the right tools.
Setting the memory to "fast timings" is good for 8-12% more tokens/second if you haven't tried yet. I miss the slightly older days of AMD when powerplay tables were unlocked and we could configure the timings and voltages manually, there's another 30% being left on the table ez
Then I clicked away and back, and now it goes Qwen second, with 58.4, to Opus Max at top with 59.2.
I have screenshots of both. The description above the chart is the same in boh cases:
> Artificial Analysis Agentic Index > Represents the weighted average of agentic capabilities benchmarks in the Artificial Analysis Intelligence Index (GDPval-AA v2, ³-Banking)
What happened? How can the scores change so much in a few seconds?
https://artificialanalysis.ai/methodology/intelligence-bench...
Edit to provide AA's article explaining it:
https://artificialanalysis.ai/articles/artificial-analysis-i...
Interesting that they chose a nano-sized model from OpenAI to be a grader for benchmarks involving knowledge and hallucination.
The order changes but I think the story discussed in this thread holds - this is a very impressive release and Qwen3.8 Max is a huge step up in agentic capabilities.
Relevant blog post (also linked to by others): https://artificialanalysis.ai/articles/artificial-analysis-i...
Do you mean the timing looks like: "We're SV tech-bros. Our benchmarks showed a chinese model above what's considered the best model at the moment. So we quickly modified the benchmark so that our SV tech-bros don't look like they're losing to a chinese model"?
That's indeed a bit fishy.
I'm very much looking forward to their forthcoming smaller model Qwen 3.8 releases. A version that can easily run locally would be great.
OpenCode or oh-my-pi might make more sense if you just want a batteries-included agent. You can also make Claude Code work with other models without too much work, but I think that's asking for headaches.
Especially the second one seems exactly like my experience.
It's a simple switch to make: cursing = try harder instead of cursing = stop trying. Is it really impossible to train Claude that way?
With weaker models you can sort of understand, they're trying their best and failing, but this thing just channels its immense inteligence into being as annoying as possible instead. I know it can do what I'm asking it to do, but it just finds a way to weasel out of it, or maybe just thinks for 10 minutes instead, then fixes one thing and breaks four additional ones.
But even fable has the annoying tendency to invent new jargon and produce an incomprehensible soup of text.
"give me this again without jargon invented this session at high density
and with a couple (maybe more or less) simple useful ascii diagrams underneath each design"
The context is that I was discussing an experimental new idea for my video game review analysis product.
Designs 1,2, and 3 were horrible: the model even suggested a rejection after the word soup so it would have been pointless to waste my fleeting time on Earth reading it.
Otherwise, I generally really enjoyed using fable for bouncing ideas. It was an absolute joy to have this thing provide useful criticism, analyse sample data, and create prototypes so that I could elevate my understanding of the problem without stepping down from a pure intuition/design headspace.
But I don't consider the purely model written code usable for a feature this important. I'll probably scrap it entirely and start from scratch with newfound understanding.
I'd open a blog with "weird things Opus did". Today it launched a swarm of cpu-hogging processes to test if the widget showing machine and I/O load is rendering nicely and correctly. The test went fine, but it was no longer able to kill those processes since they were really effectively hogging the CPU in various ways - being diligent, some of them were hogging CPU, some were murdering the SSD, some were pounding on the network adapters. Took me 30 mins to recover the machine to a working state without killing the meaningful, messy, in-flight sessions i had going on on other projects.
Infuriatingly so, in a way I don't remember Opus 4.8 being, but maybe I've just been ruined by Fable 5.
I got so used to it, when they finally pulled access for me and I had to go back to Opus I felt like I was working with my hands tied.
I finally know what those women with AI boyfriends felt like when their app updated and it won't dirty talk with them anymore.
Like the Fable ban stunt, I wouldn't put it pass Anthropic to kneecap Opus deliberately to drive more people to their more expensive option.
To me it's not so much the dumb mistakes (although there are some of those) but the ultra-verbose, mega-inefficient "solutions" to some problems / prompts.
Stuff that "works" if you're the kind of person that considers slamming a semi-trailer at 200 mph into a door did, technically, result in the door being somehow "open".
As it's supposed to be one of the most advanced model, I can't help but wonder if the solutions are that bad/verbose/inefficient because we're already in a loop of models being trained on sloppy-pasta from previous models.
“One thing worth your attention, if you were to detonate a pipe bomb in your house, it would have a negative effect on your living room”.
It will do everything it can to defer or push it off, to the point where I’ve had to add multiple imperative directives to the AGENTS file telling it, in no uncertain terms, not to defer tasks under any circumstances.
• Qwen3-VL picks up new images in a NAS, auto captions and adds the text descriptions as a hidden EXIF layer into the image, which is used for fast search and organization in conjunction with a Qdrant vector database.
• Gemma3:27b is used for personal translation work (mostly English and Chinese).
• Some small 8b models (like llama3.1) for sentiment analysis on text.
But haven't really tried using local LLMs in conjunction with agentic harnesses yet.
My secret sauce is to use LLAMAcpp's reasoning-budget and reasoning-message that trigger cut off to overthinking with a message that says to either us subagents or compress the context. opencode's dynamic context pruning plugin can get you pretty far into the stratosphere.
See https://news.ycombinator.com/item?id=48883538 25 days ago
> The Sleev (the project has been renamed to make a startup) creator was shilling their project in the OpenCode Discord. That person is very convinced they have something that no one has ever built before. They focused on token reduction without any real evals for capability impacts.
I'm generally against this context pruning without prompting or details. Sleev is very opaque about how it works and definitely will bust your cache.
Thanks for the tip - I like this a lot. I remember having to do a lot of tweaking to curtail Qwen QwQ-32b when it would go down an endless psychotic recursive reasoning loops as part of its "chain of reasoning."
I havent yet tailored the pruning messages, but mostly it works.
Reasoning budget can also be set by client, so potentially smarter.
I could not get Opus 5 to do anything without losing a few years of my life from stress.
Fable has been okay but I am doing ML work and not allowed to use it which feels insane.
I don’t see anyone talking about how you have to completely change your prompting strategies with Op. 5 versus 4.8 to get the most success.
After I started reading complaints about Opus 5, I gave Fable the task of evaluating a bunch of code Opus 4.8 had written and compare it to Opus 5's code. Fable ran a dynamic workflow and the scores came back 15-20% higher for Opus 5's code in terms of quality, correctness and readability/conciseness. I did not tell Fable which Opus wrote which code, and I turned off memory as well to ensure there was no pollution from that angle.
My only complaint is that Opus 5's prose is annoying as hell. I wrote a custom skill for it for concise debriefs and it has been working pretty well for me.
I’m doing the same right now, and I’ve found that asking for “simple English” works most of the times, although not always.
Did you find better wording that works consistently?
Opus 4.6 is the last model that's actually useful and can "adjust" its perspective to use the newer & better solution.
Where Opus 4.8-5 has over fit training on worse/older but "dominant" solutions it refuses to adjust.
Not only does this create an existential threat to adopting progress but it also means that if you have a code base that has rare but real world tradeoff the newest versions of Opus 4.7, 4.8 and 5 are worse than useless and become a major dev timesink.
Our leaderboard combines Arena ELO, AA Intelligence index, latency and speed and goes: #1 Opus 5 #2 Kimi K3 #3 Qwen3.8 Max #4 GPT 5.6 Sol
Source: http://pellmell.ai/leaderboard.
This jumps around a lot based on the top throughput and latency of whatever provider happens to be best at the moment.
Opus-5 is practically unusable (for complex tasks) in this sense - its updates are voluminous, and dense with cryptic language (there are numerous reddit threads complaining about this, so it's not just me). I often have to ask it to re-state concisely in plain terms.
For a fairly gnarly task, after fighting with with Claude-Code + Opus-5, I ported my session to Codex + GPT-5.6-sol, and it was like a breath of fresh air.
Arguably a key aspect of intelligence is concise, clear communication, and current benchmarks miss that, at least as far as I'm aware. I would think some arena-type benchmarks where humans rate responses would measure this, though I'm not sure which those are.
https://code.claude.com/docs/en/output-styles
This community is pure trash.
"You can adjust the output style in your '.claude/settings.local.json' file".
OR
"You can decrease verbosity by doing x, y and z."
Try asking them to make useful diagrams for some stuff in a codebase, out of the box without excessive hand holding they don't make good choices about what's worth communicating and how to do it.
You see this in their pointless frontend copy all the time too.
Same concept of "over-sharing" seems to prevalent in a bunch of domains when it comes to LLMs, sometimes more visible, sometimes less.
If the AI is communicating to me and can't select the appropriate jargon level, it's failing at communicating effectively.
Yeah, makes sense. A parent wouldn't bother explaining the details of their job to a toddler.
I wonder if this is a side effect of MoE models — they can write excellent prose, but not simultaneously with writing code.
And writing doesn’t have validators like code so you can’t really scale it in the same way
Yeah if you ask it a medical question, it answers in that impenetrable jargony style that clinical journals use... full of unnecessarily custom adjectives ("orthopedic" instead of "of the bone") and discipline-specific terms (anterior, distal) even when the user didn't display mastry of this terminology (hint to frontier labs: add training cases for this; it will improve your model's usability).
My theory is that LLMs perceive the writing styles of various fields as being like different (but related) languages, and they're inclined to answer a question in the language of its source material unless specifically asked otherwise. If you add "ELI5" the model treats it as a question plus a translation task.
I think this is why programming questions are answered with an exaggerated cringey form of HN-speak ("load bearing", "gate" as a verb, "dissolves") by some models.
I always get Haiku to rephrase anything human-facing.
I mean, I get it: how fast a model responds is relevant. But a test that changes second by second is far less relevant than a test that tells you how smart the model is, and accounts for latency in some way that isn't constantly changing the result.
Latency and throughput matter a ton as a user, so I think it's actually totally defensible for a leaderboard to bounce around a lot as these numbers change. The best model to use changes a lot based on these!
But: I've been very impressed by the larger Qwen Models, and a brief try of Kimi also impressed me.
A lingering sense of quality degradation when going deep remains.
But that's not an accusation: they seem to be hitting the compute/quality tradeoff extremely well.
And on-prem capability is simply irreplaceable.
Apart from all the innovations that were driven by the strive for this optimization: quantization, "distilling" (without obvious mad-cows-disease)... I think China was an invaluable player in this progress. Intuitively, I'd even go so far to speculate that LLaMa wouldn't exist without the competition.
$0.36 per task, Intelligence Index score 56 -> Grok 4.5 high
$1.13 per task, Intelligence Index score 58 -> Qwen 3.8 Max
$0.81 per task, Intelligence Index score 59 -> GPT 5.6 Sol xhigh
$1.80 per task, Intelligence Index score 63 -> Opus 5 xhigh
> Artificial Analysis Agentic Index: Represents the weighted average of agentic capabilities benchmarks in the Artificial Analysis Intelligence Index (GDPval-AA v2, Tau³-Banking)
> Artificial Analysis Coding Agent Index v1.3 incorporates 3 benchmarks: DeepSWE, Terminal-Bench v2, and SWE-Atlas-QnA
Qwen3.8 Max is 55.4 on the Agentic Index but hasn't been tested for the Coding Agent Index.
https://artificialanalysis.ai/models/qwen3-8-max
Doesn't have the claim either. Clickbait?
Even then, this seems a much more marginal win than the headline suggested to me.
I've had to use it a bit for work, and it's been remarkable watching the degradation in performance with the default suggested current models (Opus 5 as a prime example) vs the models that got them huge attention a year ago (Opus 4.6)
If you give 4.6 a spec, or existing code to implement a feature in, it will ask some pointed questions if there's something unclear in the spec, and then produce a plan and move to implement it.
5 will freak out at even a basic task, ask itself if it's own assumptions or your instructions are correct, proceed to re-assess it's own plan, and it's instructions 3-4 times, and then maybe produce code after burning several hundred thousand tokens (and quite a bit of time) analyzing existing code and thoroughly sweeping it for irrelevant problems both to the task it was given and the spec it came up with.
It's quite bizarre to me how well advertised the benchmarks and anecdotes from people one shotting MVP browser games are, compared to the experience of everyone I know that's had to actually use it to accomplish even a relatively basic task.
That vibe coding they brag about as if it was a good thing, it shows.
Take their notation for describing permissions. The docs are not comprehensive, and in practice it doesn't quite work how they describe it.
Or their management of sub-agents. I once lost a sub-agent, it finished and disappeared from UI. Apparently, you can't bring it back yourself: you have to ask the parent agent to do it for you. But the parent was Fable, and I ran out of credits, so I was locked out of using my opus sub-agent because of it.
Or an even more grotesque example: when you paste your claude API token to authorize, it covers characters with *. But it seems like an LLM has hallucinated a limit of API key length and the tail of your key stays visible.
I've probably gone to file 20 bugs. In all 20 cases there wasn't just one issue already filed for it: there were several, each which had a bunch of upvotes. And in all 20 cases ... every. last. one. ... Anthropic closed the ticket with no comment.
IF YOU ARE GOING TO HAVE A SHITTY VIBE CODED PRODUCT, AT LEAST USE YOUR SHITTY AI TO FIX THE SHITTY PROBLEMS!
I can't believe how many critical bugs fall through.
My favourite one is the bug where Plan mode can execute destructive commands inadvertently.
Then all these get closed with `Closing for now — inactive for too long. Please open a new issue if this is still relevant.`. Awesome.
Almost like CC is 100% vibe coded.
Opus 5 is just a token burner.
I use fable plan and spawn opus 4.8 workflows which seems to work alright.
Of course, the hilarious thing to me is that Anthropic likes to claim that the usage limits are because of resource allocation problems or something like that. Obviously no such issue exists, otherwise they wouldn't allow you to bypass it by just paying a bit more and it would be a hard limit. So usage credits are entirely their way of just screwing you out of more money.
For me it is not great for design work - Fable is way better, and 4.8 was conservative and thus better (Opus 5 seems to jump to conclusions far more eagerly). But for overnight builds, where I give it 8hrs worth of work on LLDs created by Fable - its great. Where Opus 4.8 would often lose the plot and stop for questions clearly answered in the LLD - Opus 5 does manage to complete. Since it launched, I don't remember it ever disappointing me with builds. But designs? Boy, is this thing explosively stupid sometimes.
With the $200 subscription, I can have Fable on ultracode working for hours and not dent the usage limits.
I don't think you're right about that last prediction, at all. And new use cases are opening up as these get smarter. I think things are going to get pretty weird.
But the point was that it really doesn't look like they're losing money on users, on average.
There is some element of responsibility on the user to guide and monitor the model/harness and not let it rip to burn tokens.
Maybe they consider that hiring a person to do it would have cost at least as much and taken much more time, so paying them is a bargain.
Plus we get to own, keep, run, do whatever with the model. We don't feel trapped. Moreover, it's something we can truly build on top of and own our own destiny.
Anthropic and OpenAI are the new Oracle (Oracle pre-AI; Oracle is even worse now). Expensive, feels like dealing with a lawyer, and not at all open. They just became infinitely less cool than they were a month ago.
The whole of our industry is going to migrate to open weights. We're smart enough to know this is the better deal and technical enough to be able to pull it off.
The only thing that might save these OpenAI and Anthropic in the near-term is an abundance of enterprise contracts negotiated with non-tech companies. They'll soak consulting firms and F500 companies for "AI" integrations.
I think that's exactly what they are going for - enterprise and government customers.
Amazon's first principle is the Customer Obsession. Making customers happy.
Fun bit is that the human psychology rates personal looking fixes better than having no issues at all.
For example, AWS overcharges you, you contact support, and more or less hassle free they refund or issue credits. The customer feels appreciated, or at least got something "extra" or "special treatment".
Meanwhile, any other (small) cloud. Simple, no weird charges. Even _most_ of network egress is free. But, no reason to call support or feel "extraordinary". Comes out as "meh" against Amazon's "top tier" support model...
Anthropic trains models on AWS's (and GCPs, and Microslop's) infrastructure, then skims margin off of selling inference also on the infrastructure owned by the other companies.
These are extremely different businesses.
People produce such models by over-RL-ing smaller models on math and coding tasks. I've found the results capable of neither innovative work nor thinking outside the box. They're straight-A students raised by tiger moments who never let them play freely for hours in the dirt.
Perhaps you could say such models are skilled --- but intelligent? Not by my measure.
People and AIs alike need diversity of experience and a broad liberal arts education to see hidden connections between fields and make real advances.
Sticking to LLMs, they seemingly get their intelligence (whatever that really means) from building models rich with knowledge, so you could have a point. But Qwen models seem to be particularly good, even at small model sizes, at maintaining both their own knowledge while acquiescing to and integrating external information in the moment.
But if you are sort of pair-programming with the model, the speed obviously matters and I think then the 35B is acceptably smart, and when it's wrong it'll be wrong much more quickly. It seems very good on SQL and PHP, and I assume on typical JS and Python.
I would rather work that way, so I hope they do produce a small MoE model.
I've been trying it on several projects and have found it's pretty sloppy. It leaves stuff broken, doesn't reliably write tests to check its own work unless explicitly prompted, misunderstands the assignment, etc.
It is smart and reasonably quick but not reliable.
At their best, I think they're closing in on Opus and GPT, but they're incredibly inconsistent and the variance in output quality is much higher than the best from any of the Anthropic or OpenAI models from the last few generations. The only way I can describe it is that it feels like a lack of intuition with the models which means I find my self needing to write longer prompts or have more back and forth to get them to do what I want from them.
To give an example, I have a saved prompt that I use as a sanity check on some data I'm storing. It reads about 50 rows from a DB and matches them to the UI and makes sure the data is displaying correctly. I've been using this with GPT 5.5 and now 5.6 for a few months and running it a few times a week with no issue. Sometimes I'll run it multiple times in a single chat if I notice bad data (run it, fix thing, run again, fix another thing).
I recently tried to switch to using Deepseek v4 (first flash and then pro) and while both did the task just fine, both would do things like change the response format from one message to another in the same chat or randomly decide to omit things it didn't think were relevant. At one point I ran the prompt, fixed some bad data, and then said "Okay, I fixed row 7, run {prompt} again" and so it decided to leave row 7 out of the response. A few times the first message would contain a table and then the next run in the same chat would contain the data in a bulleted list.
None of those are major issues and all could be solved with a bit more rigor in my prompting, but for me it makes them harder to work with. Those examples are a bit trivial, I think they're the easiest way for me to illustrate the gaps I see with them.
I wasn't able to find an explanation from them. Anyone knows what happened?
> launch traffic hit our public API endpoint harder than expected, causing intermittent instability.
https://x.com/QwenDevs/status/2085279963654275247
Many providers will host it and will compete on price. It also can't easily be taken away because one company (or one government) decides they don't want it around any more. People can fine-tune it for particular workloads.
Might as well use gpt-sol.
I stopped paying attention to self-published benchmarks when Apple started using those non-sensical performance graphs with "relative performance" as a vertical axis when announcing a new chip.
It's barely better, and barely cheaper, not really enough to challenge the status quo IMO. Half the price for basically the same performance would be a much stronger value proposition.
Things change radically month to month. Nobody is remotely close to capturing the market or having any kind of stability over time. People move around quite a lot, often to sidegrade within a generation. Just playing fly on the wall with discourse would be enough to tell you all of this, even without the data to back it up.
Also, OpenRouter misses most of the usage of the US models, as most people are getting those from the vendor directly via subscriptions.
You can argue there's a selection bias that openrouter users are less likely to display model loyalty, but it would still be a visible confounding factor if it was a statistically significant behavior. And it's not. Nor is there a visibly meaningful indication that people don't sidegrade between models. With every single data set, you're going to see that. You're also going to see it reflected in discourse, as I mentioned. Fact of the matter is there isn't a status quo in AI any more than there's a status quo in cars.
That said, it's a fair point. For me, it boils down to things covered here: https://earendil.com/posts/session-portability/
Things like obscured reasoning traces.
There are a couple of frontiers (ok bad word, maybe categories) in open weight models.
These Qwen 3.8 and Kimi K3 style models aren't trying to win on price, they're trying to compete on intelligence and capability.
Models like Deepseek V4 Flash (updated this week) are $0.03 a task, or 50X cheaper than Qwen3.8/Kimi K3, and 100X cheaper than Fable, while offering stunning intelligence. That's a different frontier for competition, and perhaps one more interesting for someone who wants to see them compete on cost.
What cost the most in API. Input, Cached Input, or Output. There you have your answer.
Unfortunately, we have moved so much of the actual intelligence of models towards reasoning, what results in some models getting good scores, but this is because they are dumping a insane amount of reasoning tokens at the problem.
So a mid priced model, with heavy reasoning output, cost the same as a expensive model, with medium reasoning output.
Before the GPT Luna price drop of 80%, you actually had the same price if you used Luna High and Sol Low. With the difference that Sol Low was insane fast, and often way better code.
https://deepswe.datacurve.ai/
Do not look at the top score but more what is on the horizontal axis as you go down. Sol Medium is frankly, was the best performance for dollar, until that Luna price drop. I will even argue that despite the higher price, Sol Medium is still way better despite Luna Max being cheaper. Or Opus Low, one of the better values also.
What do you notice? Is that those models all have a high intelligence start point for their low setting. So that means they do not rely as much on output tokens aka thinking.
Qwen3.6-27B-FP8 currently espouses (see below), which is not too bad. The directions are a bit mad but the mileage is about right and there is a helicopter manufacturer here and a RNAS (navy not airforce) museum nearby at Yeovilton. Cosford is in Shropshire which is not a million miles away.
I'm not sure what 盆地的 means but the river Yeo is correct ... OK ... "basin like" - again not bad, even if Chinese is not the first language here. The model understands that Yeovil is named after (or vice versa or at least is associated with) a river
Yeovil is the current form of Gifle (Saxon) which I thought meant "bend in a river" but WP is currently saying "fork in a river". My source is a local museum. There is a fork but was it there 2000 odd years ago? My hydrology skills say ... possibly
---------------------------------------------------- Q: where is yeovil:
Yeovil is a town in Somerset, in the South West of England.
It is located roughly:
Yeovil is known for its historic market town center, RAF Museum Cosford (nearby), and as a significant industrial town, particularly during World War II for aircraft manufacturing (including the Wellington bomber). It sits in the盆地的 valley of the River Yeo.I was with you until there. Qwen and the OpenAI models are great, aggressive agents, but they’re not as good as the anthropic models for human interaction. They just don’t have the subtlety, understanding, or attention to detail.
Opus 4.7/4.8/5? Absolutely smug and antagonistic and preachy. I'm constantly fighting with it to stop fighting me and accept that I occasionally know better. It's really frustrating to spend so many tokens of such an expensive model arguing with it.
Clearly the weighting of those things depends on the usecase
$25k - DSv4 Flash
$4k - Qwen 3.6 35A3B Q5
$1k - Qwen 3.6 27B Q4
Some people prefer the sense over the MoE YMMV.
2 DGX Sparks can run DS4 at 1 million context with 50 TPS so that's $8k.
1 A4500 can run 35A3B. Those are about $1200 new.
27B actually takes more hardware to run than 35B because attention is done differently I believe and therefore KV Cache takes a lot of space. It will run on an A4500 but it's slow and context will be like 32k.
[0]https://humanparadox.org/local-vs-frontier-benchmarks-for-my... - note here I tested Q8 but have found no difference at lower quant.