Artificial Intelligence thread

siegecrossbow

Field Marshall
Staff member
Super Moderator
Southeast Asia will be the next data center hotspot (because the Middle East is a bit of a mess, and nobody will choose Europe), where a large number of Chinese companies are already building servers using Nvidia chips; the infrastructure is already there. If the executive order is indeed issued, some hosting providers will definitely choose the move there

The current executive order is insufficient to prevent such behavior.
Problem with South Asia will be reliable power infrastructure. Of course there will be state run companies in China that would love to help the out on that regard. Hopefully BJP behaves like a moron and reject those investments on grounds of national security.
 

tamsen_ikard

Captain
Registered Member
Qwen did a wierd launch for its latest model. They did not release benchmark scores. They are also saying qwen will get better with usage.

Why would people get interested to use it if its not mature and still evolving on the fly?

That's not how you release a model. You release by giving out benchmarks to create buzz and people get interested and get media hype.
 

AI Scholar

New Member
Registered Member
Qwen did a wierd launch for its latest model. They did not release benchmark scores. They are also saying qwen will get better with usage.

Why would people get interested to use it if its not mature and still evolving on the fly?

That's not how you release a model. You release by giving out benchmarks to create buzz and people get interested and get media hype.
They’re likely rushing to release ahead of upcoming models from GLM and DeepSeek. By launching an open beta, they can gather feedback and refine the model on the fly, otherwise they’d risk being overshadowed, especially if ByteDance also drops Seed 2.5. The competition is very high between Chinese AI companies.
 

9dashline

Major
Registered Member
They’re likely rushing to release ahead of upcoming models from GLM and DeepSeek. By launching an open beta, they can gather feedback and refine the model on the fly, otherwise they’d risk being overshadowed, especially if ByteDance also drops Seed 2.5. The competition is very high between Chinese AI companies.
Also there is a rumor Xi instructed for Qwen to release its largest model... The MAX versions have never been released before
 

AI Scholar

New Member
Registered Member
Also there is a rumor Xi instructed for Qwen to release its largest model... The MAX versions have never been released before
I don't think it's just a rumor, the official Alibaba Group Twitter account retweeted this. I wonder if ByteDance will also be instructed to release its Seed 2.5 model. That would really improve the open model scene. Screenshot_20260721_001247.jpg
 

tokenanalyst

Lieutenant General
Registered Member
I don't think it's just a rumor, the official Alibaba Group Twitter account retweeted this. I wonder if ByteDance will also be instructed to release its Seed 2.5 model. That would really improve the open model scene. View attachment 178519
1784608579571.png

Why for these people every that happens is connected to Xi? It could not be that after the disbandment of the previous team Qwen models have really fallen out of favor in comparison others Chinese labs like DeepSeek, Moonshot and ZAI both in performance and usage, they are trying to regain the community back by open sourcing their biggest model.
I hope the release smaller local models as well.
 

bsdnf

Senior Member
Registered Member
super nodes are the future. Biren has told media that Chinese models are going to 5T+ parameters next year and you can see the nodes coming out to support 5 to 10T models.

Keep in mind, that Chinese models stayed smaller (Kimi had the largest ones at 1.5T until recently) due to these limitations. SuperNode, can hold all the memories.

I think the bottleneck is still in the software. Longcat 2.0 only proves that training a 1.6T model is possible within the Huawei ecosystem. However, the training efficiency and whether it can be scaled to other chip and server ecosystems remain issues.

But it will undoubtedly accelerate
 

tphuang

General
Staff member
Super Moderator
VIP Professional
Registered Member
Qwen did a wierd launch for its latest model. They did not release benchmark scores. They are also saying qwen will get better with usage.

Why would people get interested to use it if its not mature and still evolving on the fly?

That's not how you release a model. You release by giving out benchmarks to create buzz and people get interested and get media hype.
they are just updating check points until they are satisfied with it. DeepSeek had been experimenting with new V4 checkpoints also.

More domestic compute using Supernodes are coming up with different players.

 

tphuang

General
Staff member
Super Moderator
VIP Professional
Registered Member
Impressive. I believe Zhipu is the first major Chinese AI lab/company to start using only Chinese AI chips exclusively for its data center, no? They were also the first in China to start training their multimodal AI model on just Chinese chips(Huawei ascend chips) the beginning of this year if I remember correctly ?

Among the big 4, the one using the most Ascend chips right now is DeepSeek. All their post training are done with Ascend. Zai is still probably doing most of the training and research on Nvidia.
 
Top