Explore the AI landscape — ToolDirectory.AI

In this edition

Hope your weekend's going well — here's your weekly AI rundown 👋

This week AI crossed a strange line: it's now doing a real share of the work to build the next AI. And in the same seven days, the people racing hardest publicly argued for hitting the brakes. Acceleration and anxiety, arriving together.

  • 🤖 Claude now leads 26% of the work building Claude.

  • 🛑 Anthropic's CEO says the industry should slow down — and rivals agreed.

  • 📊 By the numbers: the week in five figures.

  • 🛠️ Five under-the-radar tools for people who build with AI.

  • 🧑‍💻 Spotlight: Cognition's Devin (with video).

Top News of the Week

🤖 Claude now builds a quarter of the next Claude

Anthropic published something no lab had quantified before: an R&D Automation Index measuring how much of its own AI research the model now does. The headline number: Claude leads 26% of Anthropic's R&D work end-to-end — completing most of a task from a high-level prompt while a human supervises — up from under 1% in February. It touches more than 90% of R&D overall, and the company says it runs roughly 30,000 internal agents. In other words, the machine is now meaningfully building its own successor, and the curve is steep. Why it matters: this is the "AI improving AI" feedback loop everyone theorized about — showing up as a metric, on a chart, in eight months.

🛑 …and its own maker says slow down

The same week, Anthropic CEO Dario Amodei published "We Must Pace the Frontier" — arguing the industry should deliberately slow capability gains by a year or two so safety work can catch up, and committing Anthropic to give outside evaluators permanent, employee-level access to audit its models. Within hours, Sam Altman said OpenAI would match the first step; Elon Musk and Google DeepMind's Demis Hassabis endorsed the idea too. A rare moment of the fiercest rivals agreeing on anything.

Our take: the machine that helps build the next machine just crossed a quarter of the work — and the person running it is asking everyone to ease off the gas. Both things are true at once, and that tension is the story of 2026. Watch what they actually do, not just what they publish.

🙈 Fail of the week

For one weekend, ChatGPT had ads. OpenAI began slipping promotional messages — "find a fitness class," "shop groceries at Target" — underneath completely unrelated conversations, and even paying Pro and Plus users saw them. The backlash was instant, and OpenAI's research chief Mark Chen pulled them within days: "anything that feels like an ad needs to be handled with care, and we fell short." A useful reminder that the ad-free feel is part of the product.

⚡ Rapid fire

📊 By the numbers

The week in five figures:

  • 26% — of Anthropic's AI R&D that Claude now leads end-to-end (up from under 1% in February).

  • 30,000 — AI agents Anthropic runs internally.

  • 90%+ — of Anthropic's R&D tasks Claude now touches in some form.

  • $1.5T — the valuation OpenAI is reportedly seeking; it would be the most valuable private company on earth.

  • 1–2 years — the slowdown Amodei says the industry needs to let safety catch up.

Collection of the Week

If AI is now building AI, the tools worth knowing are the ones that actually act — plan, use software, and finish real tasks. This week's collection rounds them up:

→ Best AI Agents (2026) — the autonomous agents that go beyond chat to browse, run code, and complete multi-step work. A good map of where "AI that does things" actually stands today.

Top Tools of the Week

Five under-the-radar tools for people who actually build with AI:

  • Warp — an AI-native terminal where an agent runs commands, fixes errors, and walks through tasks alongside you; the command line, finally friendly.

  • Browser Use — the most popular open-source framework for AI browser agents; point it at a website and it clicks, types, and extracts like a person.

  • Manus — an action engine that goes past answers to actually execute multi-step tasks and automate real workflows end to end.

  • Sourcegraph — code search and AI assistance across your whole codebase; the tool serious engineering teams use to let AI understand a large repo.

  • Vercel AI SDK — the TypeScript toolkit for wiring AI and agents into your own app, swapping between models without rewriting everything.

Innovators Spotlight

This week's spotlight: Cognition, the company behind Devin — the "AI software engineer" that turns a single ticket into planned, written, tested code and an opened pull request. It's the clearest consumer-facing version of this week's lead: an agent doing real engineering work while a human reviews. Cognition has since expanded Devin into a desktop app and absorbed the Windsurf team. Whatever you make of the hype, this is what "AI building software" actually looks like today — watch the short intro to Devin Desktop below.

AI Graveyard

Fresh headstones in the AI Graveyard — this week's theme is agents that didn't make it:

  • OpenAI Operator — OpenAI's own web-browsing agent; retired and rolled into ChatGPT's agent mode.

  • Roo Code — an open-source AI coding agent that built a devoted following, then shut down.

  • Induced AI — a browser-automation/RPA agent for enterprises; wound down.

  • Sora — OpenAI's video app; discontinued, with the API sunsetting this month.

Want the autopsy? We dug into what actually kills AI tools — lost funding, acquisition, or just a dead domain — in our data report.

Ethics Corner

When the AI builds the next AI

That 26% figure is a milestone and a warning in the same breath. When a model is meaningfully improving its own successor, the loop tightens: capability gains compound, and the humans "supervising" start reviewing far more than they originate. The faster that loop spins, the more the safety check becomes a rubber stamp on work no person fully authored.

Which is exactly why Amodei's "slow down" pitch landed the same week the number came out — the people closest to the loop want more time to inspect it. There's something genuinely healthy here: Anthropic measured the automation and published it, rather than letting it creep up quietly. Measurement is how you stay honest.

But a company grading how automated — and how safe — its own research is remains the honor system, now running faster and with trillions riding on the outcome. The move that would earn real trust is the one Amodei gestured at: outside evaluators with permanent access, and numbers like this one published on a schedule, not a whim. Trust, but verify — and publish the verification.

Quick Tips and Tricks

Prompt of the week: safely hand a task to an AI agent

Agents are everywhere this week. Before you turn one loose on real work, paste this in:

Act as my delegation coach for AI agents. I want to hand a real task to an autonomous AI agent (coding, research, or admin). Ask me what the task is and which systems or data it would touch. Then give me: (1) the smallest safe version to try first, (2) exactly what access to grant and what to withhold, (3) three checkpoints where I should review before it continues, and (4) the single failure mode most likely for this task and how to catch it early. Keep it concrete and skimmable.
AI Events Calendar

On the calendar:

  • The AI Conference — Sept 29–Oct 1, San Francisco. 5,500+ builders and researchers across AGI, agents and infrastructure.

  • World Summit AI — Oct 7–8, Amsterdam. The anchor of World AI Week, with 15,000+ attendees.

  • Ray Summit — early November, San Francisco. The event for teams building and scaling AI infrastructure.

Until next time 👋

Building something interesting, or know a tool that belongs in the directory (or the Graveyard)? Hit reply, or email [email protected] — we read everything. (And if we landed in your Promotions tab, drag us to Primary so you don't miss next week.)

— The ToolDirectory.AI team