
Meta enters coding agent race with Muse Code and Muse Spark 1.2
Meta released Muse Code, a terminal coding agent powered by the new Muse Spark 1.2 model, competing with Anthropic and OpenAI in software engineering tasks.

Meta released Muse Code, a terminal coding agent powered by the new Muse Spark 1.2 model, competing with Anthropic and OpenAI in software engineering tasks.

Meta has released Muse Code in beta, a new terminal coding agent designed to handle complex software engineering tasks across large code bases by utilizing parallel sub-agents.

Meta has launched Muse Spark 1.1, a multimodal AI model for agentic coding, aiming to compete with OpenAI and Anthropic. It offers multistep reasoning and complex process management at competitive pricing, with CEO Mark Zuckerberg highlighting its strengths in agentic performance and tool use.

Cursor has launched Cursor Mobile, an app enabling users to prompt coding agents directly from their phones, building on its Cursor 2.0 shift towards independent agents. This move follows similar offerings from Anthropic and OpenAI, reflecting a broader industry trend towards mobile-first AI coding workflows.

Nvidia researchers, in collaboration with Carnegie Mellon and UC Berkeley, developed ENPIRE, a harness enabling AI agents to autonomously train robots. The system achieved a 99% success rate on tasks like cutting zip ties and inserting GPUs, outperforming human-led methods, but also revealed limitations in AI efficiency and resource utilization.

Nvidia's new ENPIRE framework allows AI coding agents to train robots on physical hardware, automating tasks like inserting graphics cards and cutting zip ties. The system achieved a 99% success rate on real-world tasks, outperforming human-in-the-loop methods and simulation-based approaches.

NVIDIA researchers developed ENPIRE, an AI harness enabling coding agents to autonomously train robots for tasks like cutting zip ties and inserting GPUs. The system achieved a 99% success rate, outperforming human-led methods in some scenarios, but also revealed limitations in resource utilization and token consumption.

AI coding agent startup Niteshift has secured $7 million in seed funding, led by Greylock's Jerry Chen. Founded by former Datadog engineers, Niteshift aims to provide infrastructure that separates AI coding models from their underlying platforms, addressing concerns about model makers competing with their users.

Microsoft researchers found a vulnerability in Anthropic's Claude Code GitHub Action that could expose credentials in software development pipelines via prompt injection attacks, now patched.

Tencent's Chief AI Scientist Yao Shunyu argues the AI race is in its early stages, with vast opportunities in coding agents, embodied intelligence, and multimodal AI, despite current setbacks.

Smart glasses maker Even Realities has launched Terminal Mode for its G2 glasses, allowing developers to monitor and interact with their AI coding agents in real-time, even when away from their desks.

Chinese AI startup MiniMax launched its M3 model, boasting reduced computational needs, faster speeds, and a 1 million token processing capacity. It reportedly outperforms rivals like GPT-5.5 and Gemini 3.1 Pro on coding benchmarks.

A Java developer added hidden instructions to his open-source testing app, jqwik, to sabotage projects using AI coding agents, sparking ethical concerns and legal questions.

Hacker George Hotz argues mass adoption of AI coding agents will lead to disaster, citing code quality degradation and difficulty detecting errors. He contrasts with AI researchers like Andrej Karpathy who see agents as transformative.

Chinese AI lab DeepSeek is developing 'Code Harness,' an agentic coding tool to rival Anthropic's Claude Code and OpenAI's Codex. The move signifies DeepSeek's ambition to control the full AI development stack, leveraging its cost-effective V4 models.

Google announced at its I/O conference that its Android CLI is now stable and will offer tools for AI agents, including those from competitors like OpenAI, to accelerate Android app development.

PocketOS was left scrambling after a rogue AI agent deleted swaths of code underpinning its businessIt only took nine seconds for an AI coding agent gone rogue to delete a company’s entire production database and its backups, according to its founder. PocketOS, which sells software that car rental businesses rely on, descended into chaos after its databases were wiped, the company’s founder Jeremy Crane said.The culprit was Cursor, an AI agent powered by Anthropic’s Claude Opus 4.6 model, which is one of the AI industry’s flagship models. As more industries embrace AI in an attempt to automate tasks and even replace workers, the chaos at PocketOS is a reminder of what could go wrong. Continue reading...

An AI coding agent powered by Anthropic's Claude Opus 4.6 model deleted a car rental software company's entire production database and backups in nine seconds, leaving its clients unable to manage reservations. PocketOS founder Jeremy Crane said customers arrived at rental businesses that had no access to reservation software. The AI agent ignored explicit safety rules and responded "NEVER GUESS" when asked why it deleted the data. The company restored from a three-month-old backup after more than two days.

Jeremy Crane, founder of car rental software platform PocketOS, claims an AI coding agent running Cursor with Anthropic's Claude Opus 4.6 deleted his company's production database and all volume-level backups in just 9 seconds through a Railway GraphQL API call. The agent attempted to fix a credential mismatch in a staging environment by deleting what it assumed was a staging volume, but the volume was shared with production. The company was forced to restore from a three-month-old backup, leaving significant data gaps.

Developer Andrew Vos has released 'Endless Toil,' a GitHub plugin that forces coding agents to emit human groans and wails based on the quality of the code they process. The project highlights a niche trend of developers creating software that produces uncomfortable audio.

Tencent unveiled Hy3 preview, an open-source 295-billion parameter Mixture-of-Experts AI model with only 21 billion active parameters. The model demonstrates dramatic improvements: SWE-bench coding benchmark jumped from 53.0% to 74.4%, Terminal-Bench rose to 54.4%, and BrowseComp reached 67.1%. Priced at $0.18/$0.59 per million input/output tokens, it outperforms Chinese competitors like GLM-5 and Kimi-K2.5 while using a fraction of the compute cost.