AI Industry Daily, 20 Sep 2026

1. DeepSeek updates its peak/off-peak API pricing, treating make-up work weekends and public holidays as off-peak throughout; 2. StepFun's Step 5 Preview appears on the Artificial Analysis leaderboard with million-token context and vision; 3. A wave of open-source releases from Chinese institutions: DAMO Academy's abdominal CT model RADAR published in Science, China Telecom open-sources the Xing4.0-29B-A4B MoE model trained on domestic compute, and Qwen releases Qwen3.8-LiveTranslate for real-time interpretation across 60 languages; 4. OpenAI, Anthropic, Google and SpaceXAI face a class-action antitrust suit from paying users; 5. Leaks point to next-generation models from several major vendors, including GPT-6 Sol/Luna, Gemini 4 Pro, MiniMax M3.1 and Kimi K3.1

Traduït de l’original en xinès.

1. DeepSeek updates its peak/off-peak API pricing, treating make-up work weekends and public holidays as off-peak throughout;

2. StepFun's Step 5 Preview appears on the Artificial Analysis leaderboard with million-token context and vision;

3. A wave of open-source releases from Chinese institutions: DAMO Academy's abdominal CT model RADAR published in Science, China Telecom open-sources the Xing4.0-29B-A4B MoE model trained on domestic compute, and Qwen releases Qwen3.8-LiveTranslate for real-time interpretation across 60 languages;

4. OpenAI, Anthropic, Google and SpaceXAI face a class-action antitrust suit from paying users;

5. Leaks point to next-generation models from several major vendors, including GPT-6 Sol/Luna, Gemini 4 Pro, MiniMax M3.1 and Kimi K3.1.

I. Model releases and open source

1. StepFun's Step 5 Preview surfaces on the Artificial Analysis leaderboard, with limited user access

Artificial Analysis has listed Step 5 Preview and completed independent third-party evaluation, and some Step Plan subscribers can now call the model. The leaderboard shows an intelligence index of 44, a 1M context window and vision input, priced at $1 per million input tokens and $2.70 per million output. It is currently marked as a closed model; StepFun has not held an official launch, and the model's full capabilities and general availability date are still open.

Evaluation page: https://artificialanalysis.ai/models/step-5

2. Alibaba's Qwen releases Qwen3.8-LiveTranslate for real-time interpretation

Qwen has released Qwen3.8-LiveTranslate, a simultaneous interpretation model supporting real-time translation across 60 languages. It uses an interleaved architecture with a Thinker-Talker two-module design and reuses audio cache, cutting per-character latency from 2.8s in the previous generation to 2.3s. Three new core capabilities: real-time speaker separation, source and translation output in the same frame, and long-context disambiguation. On multi-speaker long audio and multilingual interpretation evaluations it outperforms both its predecessor and mainstream interpretation systems. Online demo and API access are open now.

Announcement: https://mp.weixin.qq.com/s/Rc3CKdHAtA_NdVRN2RLuAg

3. DAMO Academy open-sources abdominal CT model RADAR, with the paper published in Science

Alibaba DAMO Academy has open-sourced RADAR (Rapid Abdominal Diagnosis with AI and Radiology), a multimodal model for abdominal CT, with the accompanying research published in Science. It was trained on 424,911 contrast-enhanced abdominal CT studies and a dataset of 15 million anatomy-level image-text pairs, without manual annotation, covering 18 anatomical structures and 146 imaging findings. AUC reached 0.913 on a real clinical cohort and 0.874–0.912 across eight external hospitals; when radiologists read alongside RADAR, diagnostic sensitivity rose by roughly 10%. Weights, code and training documentation are all open under Apache 2.0.

GitHub: https://github.com/alibaba-damo-academy/damo-radar Hugging Face weights: https://huggingface.co/radar-generalist/RADAR Paper: https://www.science.org/doi/10.1126/science.aec6129

4. Cua open-sources cua-s1-forms, a lightweight model for form filling

Cua has open-sourced cua-s1-forms, a small model purpose-built for GUI form automation: roughly 700,000 parameters with a weights file of just 2.8 MB, deployable locally. In a single forward pass it scores four actions — fill, check, click, skip — against page form elements, handing execution to the Cua Driver. Accuracy was 99.95% on a synthetic dataset and 100% on a small real-world demo set. The project is an early research prototype rather than a general-purpose model, and has not been validated against production forms at scale.

Hugging Face: https://huggingface.co/cua-ai/cua-s1-forms GitHub: https://github.com/trycua/cua/tree/main/libs/cua-s1

5. China Telecom open-sources the Xing4.0-29B-A4B MoE model

China Telecom's AI arm has open-sourced the Xing4.0-29B-A4B semantic model, with 29B total parameters and 4B active. Optimized for engineering tasks and agent scenarios, it uses mHC, MLA and MTP architecture, natively supports 256K context extensible to 512K, and was trained entirely on an Ascend 910C domestic compute cluster with deep adaptation to MindSpore and MindFormers, improving training throughput roughly 96% over the native baseline. Base, FP8-quantized and GGUF builds are all available on Hugging Face and ModelScope, supporting local inference, fine-tuning and production deployment.

GitHub: https://github.com/XingChen-AGI/Xing4.0-29B-A4B Hugging Face: https://huggingface.co/XingChen-AGI/Xing4.0-29B-A4B

II. Developer ecosystem and tooling

1. DeepSeek updates peak/off-peak API pricing: make-up weekends and public holidays all bill as off-peak

DeepSeek has published a pricing notice clarifying how billing periods are determined: Chinese public holidays, and weekends worked as holiday make-up days, bill at the off-peak rate for the full day. Even when a weekend is a working day due to a holiday shift, it does not switch to peak pricing. The rule applies to all users of DeepSeek's official API. Developers can plan batch inference and offline jobs around it to cut compute costs over holidays.

DeepSeek platform: https://platform.deepseek.com/

2. Android Developers releases the Android Bench 2.0 benchmark

Google's Android team has introduced Android Bench 2.0 to evaluate the engineering capability of AI agent and model combinations. Tasks cover complete real Android development workflows: building an app from scratch, adding features iteratively, and migrating code across platforms, scored on completion rate, UI visual fidelity, regression bugs and call cost. The new version includes several mainstream models: Gemini 3.8 Flash, Gemini 3.7 Flash, GPT-6 Astra, Fable 5.1, Kimi K3 and Qwen3.8 Max. Methodology and comparative results are public.

Docs: https://developer.android.com/bench Announcement: https://x.com/AndroidDev/status/2100622197253398669

3. OpenAI Codex lead Tibo responds to community requests for a banked reset

After the expected Codex update did not land this week, community users asked Codex lead Tibo for a banked reset. His reply, verbatim: "OK fine. But it's also still coming in Tuesday". Two readings emerged: some developers took it as agreement to grant the reset; others noted the reply never explicitly confirms one, and the only certain information is that the feature ships Tuesday — whether the banked reset applies remains officially unsettled.

Original post: https://x.com/thsottiaux/status/2101352781219258527

III. Products shipping

No standalone product launches in this edition

IV. Technical insight and frontier research

1. Medical multimodal models reach clinical reading support: RADAR shows AI can lift radiologist detection rates

DAMO Academy's RADAR work demonstrates that a vision-language model trained on large-scale clinical CT reports can screen multi-organ abdominal imaging without fine-grained manual lesion annotation. In a human-AI collaborative reading setup, the model lifted overall diagnostic sensitivity by roughly 10% for radiologists. The implication is that the next phase for medical AI is not benchmark scores but fitting into collaborative clinical workflows — though real-world multi-centre variation, medical compliance and heterogeneous equipment remain obstacles to deployment at scale.

2. A route for lightweight specialized GUI models: small parameter counts for vertical automation

cua-s1-forms recognizes form elements and decides actions with only 700,000 parameters, representing a different paradigm: rather than using a general-purpose large model for GUI automation, score interface element actions with a very small specialized model. The upside is local low latency and low cost; the limit is weak generalization — it only fits forms within its training distribution and does not transfer easily to the heterogeneous pages of the open web. The approach suits standardized form automation on corporate intranets.

3. Agent evaluation shifts to complete engineering tasks, with Android Bench 2.0 offering a new reference model

Traditional LLM evaluation concentrated on single-turn Q&A and code snippet generation. Android Bench 2.0 raises that to multi-day, multi-step complete engineering projects, and evaluates the agent plus base model as a combination. This kind of evaluation maps more closely to how enterprises actually judge value when buying, and enterprise selection will gradually shift from leaderboard scores toward end-to-end performance on complex tasks.

V. Industry moves

1. OpenAI, Anthropic, Google and SpaceXAI face a class-action antitrust suit from paying users

Paying subscribers have filed a class action in the Northern District of California naming OpenAI, Anthropic, SpaceXAI and Google. The plaintiffs allege that executives at the four companies publicly called for coordinated control over the pace of AI development, which amounts to colluding to restrain product iteration in violation of the Sherman Act — users paid subscription fees while performance improvements were slowed artificially. They are seeking class certification, injunctive relief and a finding of violation. None of the four companies has responded publicly.

VI. Outlook and market rumours

Everything in this section is industry chatter and community speculation with no official confirmation. Treat it as industry colour, not a basis for investment or procurement decisions

1. Anthropic reportedly preparing a new model to answer GPT-6 Astra in the enterprise market

Reuters, citing three people familiar with the matter, reports that Anthropic is considering releasing a next-generation model to counter Astra's momentum with enterprise customers; safety evaluation work is still under way and no release date or version has been settled. Community leakers say new Fable, Opus and Sonnet builds are in quiet limited testing on a small number of accounts. Anthropic also needs to balance model development spend against profitability targets as it prepares for an eventual IPO.

2. OpenAI may ship GPT-6 Sol and GPT-6 Luna next week, with DevDay on 29 September

Tibo and Sam Altman have signalled that originally planned content is being pushed to next week for priority release, with the remainder rolling out through DevDay; Tibo said there is enough in the pipeline for three developer conferences. Asked by community users to ship GPT-6 Sol and GPT-6 Luna, Tibo replied "What else do you want" — read by the market as those two models being candidates. OpenAI DevDay takes place 29 September local time.

Original post: https://x.com/thsottiaux/status/2101157729037586694

3. Google's Gemini 4 Pro said to be in quiet testing on Arena

Several third-party testers report that on the Gemini Arena platform, calls to gemini-3.8-flash are randomly routed to an internal Gemini 4 Pro checkpoint codenamed Argon, with noticeably better SVG drawing, 3D scenes, and web and game generation. A widely shared benchmark image was identified as GPT-Image output and carries no evidentiary weight. Google has not confirmed the Gemini 4 Pro designation, parameters, pricing or release timing.

4. MiniMax-M3.1 traces found in open-source CLI code

Community members searching MiniMax's open-source MiniMax Code CLI repository found MiniMax-M3.1 identifiers across several config files, including context window tiers of 450K/512K/1M and multiple output length and thinking effort settings. Screenshots suggest insiders have hinted the model is close to release. It is not yet wired up in the tool and its full capabilities are unknown.

Code search: https://github.com/search?q=repo%3AMiniMax-AI%2Fminimax-code+minimax-m3.1&type=code

5. Moonshot AI teases Kimi K3.1, then deletes the post

Kimi's verified institutional account on Zhihu posted a string of digits of pi with "3.1" missing from the opening, read by the industry as a hint at a Kimi K3.1 release. The post made no direct mention of version, parameters or timing, and has since been deleted. There is no official confirmation.

VII. Claw roundup

1. Compute pricing gets more granular, with peak/off-peak billing becoming a differentiator for API vendors

DeepSeek's revision — folding make-up working days and public holidays into the low-price off-peak window — reflects how granular operations have become among Chinese cloud API vendors. Enterprise customers can schedule non-real-time work such as batch inference, offline RAG and model evaluation into off-peak windows for meaningful savings. Other Chinese model API vendors will likely follow with their own refined time-window rules, and scheduling around compute cost is becoming a required skill for AI engineering teams.

2. Medical multimodal models make an academic breakthrough, but commercial deployment still faces heavy regulation

RADAR appearing in Science shows that Chinese medical large models have reached the international frontier in imaging-assisted diagnosis. But medical AI cannot reach deployment on papers and open weights alone: medical device certification, in-hospital data compliance, liability allocation and multi-centre real-world validation are the real bottlenecks. Open weights mostly serve research and secondary study; wiring a model directly into clinical diagnostic systems still requires strict regulatory approval.

3. Major model vendors enter a dense release window as enterprise competition intensifies

Taken together, the various reports suggest OpenAI, Anthropic, MiniMax, Moonshot AI and Google are all preparing next-generation foundation models, converging on enterprise agents, long context and complex reasoning. Astra's rapid capture of enterprise token share is forcing competitors to accelerate. Enterprise evaluation cycles are shortening and version churn will keep speeding up, so procurement strategy should leave architectural room to swap base models quickly.

VIII. Trending open source on GitHub

Trending AI on 20 Sep 2026

1. alibaba-damo-academy/damo-radar

Why it's trending: medical multimodal model, stars climbing fast today DAMO Academy's abdominal CT diagnosis model RADAR, the open-source repository accompanying the Science paper, with full data processing, training and inference code under Apache 2.0, aimed at medical imaging AI research.

Repository: https://github.com/alibaba-damo-academy/damo-radar

2. XingChen-AGI/Xing4.0-29B-A4B

Why it's trending: MoE model trained on domestic compute, attention from Chinese developers spiking China Telecom's open-source Xing4.0-29B-A4B, trained on an Ascend cluster, with native 256K context, aimed at agent engineering tasks, available in FP8 and GGUF quantized builds.

Repository: https://github.com/XingChen-AGI/Xing4.0-29B-A4B

3. trycua/cua

Why it's trending: GUI agent automation project, the main repository behind cua-s1-forms A computer-use agent framework, now with a lightweight form-filling model that recognizes page elements and fills forms locally — suited to intranet automation.

Repository: https://github.com/trycua/cua

4. MiniMax-AI/minimax-code

Why it's trending: MiniMax's code CLI, where the community found M3.1 config traces MiniMax's official open-source code agent command line tool for reading, writing and investigating code, whose files now contain config fields referencing the next-generation M3.1 model.

Repository: https://github.com/MiniMax-AI/minimax-code

Tots els resums

Continua explorant què pot fer la IANavega per totes les eines