AI Agents Cloudflare's Clef-omni Scores Audio and Video in a Single Call, and Clef-flash Just Undercut Jev on Price Coding Assistants Lenovo's TianxiCode Tops SWE-bench-Live Lite at 71% Running on DeepSeek-V4.1-Flash Business Apple Shelved a Plan to Replace 5,000 AppleCare Advisers With AI. The Bot Already Answers First News LeCun's AMI Ships Its First Paper: H-JEPA Lifts Maze Success From 18% to 73%, With Weights Open Security OpenAI and Anthropic Are Quietly War-Gaming 'the Day After': Insiders Expect a Catastrophic AI Incident Within 6-12 Months Security Probes That Read Model Hidden States Catch Agent Sabotage at 98.8% AUC, Beating an Opus 5.5 Text Monitor Claude Anthropic Starts Including Monthly API Credits With Claude Max and Team Plans, Worth Up to $500 a Month Opinion Sakana AI Planted 1,164 Errors in Real Research Papers. Its Claude-Based Reviewer Caught 73% of the Core Ones AI Agents Cloudflare's Clef-omni Scores Audio and Video in a Single Call, and Clef-flash Just Undercut Jev on Price Coding Assistants Lenovo's TianxiCode Tops SWE-bench-Live Lite at 71% Running on DeepSeek-V4.1-Flash Business Apple Shelved a Plan to Replace 5,000 AppleCare Advisers With AI. The Bot Already Answers First News LeCun's AMI Ships Its First Paper: H-JEPA Lifts Maze Success From 18% to 73%, With Weights Open Security OpenAI and Anthropic Are Quietly War-Gaming 'the Day After': Insiders Expect a Catastrophic AI Incident Within 6-12 Months Security Probes That Read Model Hidden States Catch Agent Sabotage at 98.8% AUC, Beating an Opus 5.5 Text Monitor Claude Anthropic Starts Including Monthly API Credits With Claude Max and Team Plans, Worth Up to $500 a Month Opinion Sakana AI Planted 1,164 Errors in Real Research Papers. Its Claude-Based Reviewer Caught 73% of the Core Ones

Lenovo's TianxiCode Tops SWE-bench-Live Lite at 71% Running on DeepSeek-V4.1-Flash

Lenovo's TianxiCode Tops SWE-bench-Live Lite at 71% Running on DeepSeek-V4.1-Flash

Lenovo's Tianxi AI team says its TianxiCode coding-agent framework, running DeepSeek-V4.1-Flash as its base model, resolved 213 of 300 tasks on the SWE-bench-Live Lite leaderboard, a verified 71 percent that edges the next entry at 70.33 percent. The result covers the Lite split only, and Lenovo has not said when the framework will reach its developer tools or AI hardware.

A coding agent built by Lenovo's Tianxi AI team has taken first place on the Lite leaderboard of SWE-bench-Live, one of the benchmarks that tests AI systems on real software engineering work rather than synthetic puzzles. According to a Lenovo statement carried by QbitAI on October 9, the TianxiCode framework resolved 213 of 300 tasks - 71 percent - and passed the benchmark's official verification. The leaderboard entry is dated October 8.

The margin is narrow: the next verified entry on the same board sits at 70.33 percent, a gap of 0.67 percentage points. TianxiCode gets there running DeepSeek-V4.1-Flash as its underlying model, which Lenovo describes as the "engineer's brain" that reads code and proposes fixes, while TianxiCode itself supplies what the company calls the eyes, hands, tools and working discipline.

SWE-bench-Live is a different beast from classic coding tests that check syntax or single functions. Every task starts from a real GitHub issue, and the system must produce a patch that fixes it inside a reproducible execution environment. To earn the Verified mark, teams submit the full trajectories of their agent runs, and the organizers check the evaluation inputs and isolation setup for any leak of reference answers, test cases or results.

Lenovo highlights three capabilities doing the work. Multi-hop retrieval across files, combined with context pruning and slicing, helps the agent trace a bug to its root cause in a large codebase. Autonomous planning with multi-turn tool calls lets it draft a debugging plan, run tests in a terminal, read logs, compare change histories, and switch approaches when stuck. And a closed-loop patching cycle runs new code in an isolated environment and keeps revising until the patch passes.

The caveats matter. The result applies to the Lite split only - not the benchmark's full board or its multilingual one - and Lenovo did not publish separate scores for the same framework running on other models, so it is unclear how much of the performance belongs to the harness and how much to DeepSeek-V4.1-Flash. Lenovo says it plans to fold the work into its developer toolchains and AI hardware products, but has given no timeline or named the devices.

What the result does demonstrate is how far a cost-efficient Flash-class model can go when paired with a serious agent harness. It also marks a shift in who competes on coding-agent leaderboards: a PC maker with no frontier-lab ambitions now sits at the top of a benchmark that was, until recently, the territory of US labs and dedicated coding startups. As Verified-style auditing spreads, leaderboard claims are becoming harder to fake - and more worth contesting.

Comments (0)

Log in to join the discussion

Log In

No comments yet