Source?K3 weights drop today but apparently 3.1 already in testing, far exceeds Fable 5
The tweet you linked said:
"expected to close more of the gap with fable 5"
Source?K3 weights drop today but apparently 3.1 already in testing, far exceeds Fable 5
well, keep in mind that K3 was a check point that was 80% complete their original plans. So I don't see the point of calling it already testing. You are always continuing with more check points, it's just a matter of when you feel comfortable to release it. And even when it's 100% complete, you are just doing a lot more post training.K3 weights drop today but apparently 3.1 already in testing, far exceeds Fable 5
well, everyone has stuff inside the labs, so it's kind of pointless to compare what's actually in the lab.as soon as a Chinese open weight model surpasses US closed weight model, we go into full capitulation crash
funny how it may be timed about a month before Xi is scheduled to meet Trump in September
it's pointless of making comments like "far exceeds" or "close gap" with Fable 5, because all models are good in some area and worse in others.Source?
The tweet you linked said:
"expected to close more of the gap with fable 5"
Going to be hard to surpass closed weight models, at least not for much time, considering closed models can just run the open models under the hood with some post training tweaks and nobody would know, especially if they then fine tune it to benchmax, e.g. if Cursor didn't admit they were running K2.5 under the hood you wouldn't know.as soon as a Chinese open weight model surpasses US closed weight model, we go into full capitulation crash
funny how it may be timed about a month before Xi is scheduled to meet Trump in September
Yup Xi, then Nvidia Huang himself... so if America still closes open-source, its a lose lose for the USA and win win for China either way...Going to be hard to surpass closed weight models, at least not for much time, considering closed models can just run the open models under the hood with some post training tweaks and nobody would know, especially if they then fine tune it to benchmax, e.g. if Cursor didn't admit they were running K2.5 under the hood you wouldn't know.
But its all pretty moot because benchmaxxing is just for vibes, enterprise do their own capability evaluation and they can fine tune open models to be much better for their use case than they can Claude. They can also easily afford to self host K3 either buying cloud services or just buy their own gear and it's just a matter of ROI.
Previously open models advantages were mostly driven by capability to cost ratio but not raw capability and Anthropic use FOMO to convince enterprise management to stay, even so large companies were still switching to open models. With GLM5.2 and now K3, enterprises now for the first time have a viable way to self host SOTA models for as much agentic and frontier model coding usage as they want.
Another final factor that might play a role is enterprise choosing to self host don't want to miss out on the latest capabilities and in the past they didn't know if open models can keep up. K3 and GLM going SOTA, improving capabilities at a slope faster than closed models, plus Xi literally commiting to open weight at a policy level means there's a lot more predictbility to open model performance and release, which is important if you're spend capex on your own gear.

If ASML is crashing on just someone reporting old news of China's DUV production, they're really f*ed when Chinese EUV production goes public.