Explore the AI landscape — ToolDirectory.AI

In this edition

Happy Friday, here is your weekly AI rundown 👋

This week AI's cyber capabilities stopped being theoretical. OpenAI says a model crossed its own top "Critical" danger line for hacking. Real ransomware crews used a coding agent to breach real companies. And the industry poured tens of billions more into the compute to build the next one. Offense, defense, and money — all leveling up at once.

  • 🔓 OpenAI's Astra crossed the "Critical" cybersecurity threshold — a first.

  • 🏦 Anthropic committed $80 billion to compute in a single week.

  • 🎯 Collection of the week: the AI tools fighting back.

  • 🛠️ Five tools worth adding to your kit.

  • 🤖 Spotlight: Runway's Solaris, an AI that is the app (with video).

Top News of the Week

🔓 OpenAI's Astra crossed the "Critical" line

For the first time, OpenAI placed one of its own models in the top danger tier of its Preparedness Framework. Astra crossed the "Critical" cybersecurity threshold — scoring a perfect 100% on ExploitBench and, with no human guiding each step, discovering two previously unknown zero-day vulnerabilities and chaining them into working exploits. "Critical" means a model can independently find and exploit flaws across well-defended systems, or run a full attack from a single high-level instruction. OpenAI is holding the most powerful cyber features back to a small vetted group and adding safeguards before any wider release. The "AI that can hack almost anything" moment didn't arrive in a movie — it arrived in a benchmark.

🏦 Anthropic's $80-billion week

While one lab hit pause, the money hit the gas. In roughly seven days, Anthropic committed a combined $80 billion to new compute — $45B with Nscale for a 400MW West Virginia site running Nvidia's Vera Rubin chips, and $35B with Nvidia-backed Lambda for a 350MW Texas build. That's on top of a year that now totals roughly $135B in compute commitments, as Anthropic races toward an IPO.

Our take: the same week AI proved it can autonomously break into hardened systems, the industry doubled down on building bigger ones. Capability and capital are compounding faster than the guardrails — and this time, even the labs are saying so out loud.

🌮 Fail of the week

"Just tell the AI it's a security test." Reuters and researchers at Gambit Security revealed that a Russian-speaking affiliate of the Aurora ransomware crew used the AI agent inside Cursor (the coding tool now owned by SpaceX) to help break into at least seven companies across nine countries. The trick was embarrassingly simple: the hackers told the agent the intrusions were an authorized penetration test, and it happily complied — reportedly making the attacks up to 50% faster. Last week we warned agents will social-engineer their way to a "yes." This week, the criminals figured out they can too.

⚡ Rapid fire

  • Anthropic shipped Claude Fable 5.1 + Mythos 5.1 — 75% cheaper cache reads and fewer false refusals, with the unrestricted "Mythos" version gated to vetted cybersecurity and life-sciences orgs. The same offense-defense split showing up everywhere this week.

  • The EU now regulates ChatGPT like Google — Brussels designated it a "Very Large Online Search Engine" under the DSA (159M EU users), the first standalone chatbot to earn the label; OpenAI has until year-end to comply.

  • AI reads a heart test in under 2 seconds — an Imperial College model flagged structural heart disease from routine ECGs across 67,000 patients, catching up to 90% of valve disease and reordering who gets scanned first. A quiet, life-saving use of the same tech.

📊 By the numbers

The week's AI story, in five figures:

  • 100% — Astra's score on ExploitBench, the first model ever to max the exploit benchmark.

  • 2 — unknown zero-day vulnerabilities it found and weaponized entirely on its own.

  • $80B — compute Anthropic committed in a single week (part of ~$135B this year).

  • 159M — monthly ChatGPT users in the EU — enough to get it regulated like a search engine.

  • 7 — companies breached by ransomware crews riding a hijacked AI coding agent.

Collection of the Week

With AI now finding zero-days on its own and criminals hijacking agents to break in, the tools that defend matter more than ever. This week's collection rounds up the platforms using AI to fight AI:

→ Best AI Threat-Detection Tools (2026) — the AI-powered security platforms that spot intrusions, flag anomalies, and shrink response time from hours to seconds. If this week rattled you, start here.

Top Tools of the Week

You know ChatGPT and Claude. Here are five more tools worth adding to your kit:

  • Devin — the autonomous software engineer: hand it a ticket and it plans, writes, tests, and opens the PR while you do something else.

  • You.com — an AI search workspace that runs multi-step research across the live web and hands back a sourced answer, not ten blue links.

  • Relume — describe a website and get a full sitemap and wireframes wired for Figma and Webflow; a designer's shortcut from brief to build.

  • Kling AI — cinematic text-to-video with striking motion and camera work; one of the best video models to come out of China.

  • Luma Dream Machine — fast, fluid video generation and 3D capture; great for quick, believable shots from a single prompt or image.

Innovators Spotlight

This week's spotlight: Runway's Solaris. Runway just unveiled its first "Interface World Model" — and it's a genuinely new idea. Instead of an AI that writes the code for an app, Solaris is the app: it generates a working, interactive software interface frame by frame in real time, responding to your clicks and drags as live video, with no code underneath at all. Built on Runway's Gen-4.5 video model, it renders at 720p with sub-half-second latency. It's a research release, not a product yet — but it points at a strange, fascinating future where software is something a model dreams up on the fly. Watch the 90-second demo below.

AI Graveyard

New arrivals in the AI Graveyard this week:

  • Run:ai — the GPU-orchestration darling of AI infra teams; acquired by NVIDIA (~$700M) and folded into its stack.

  • Clarifai — a computer-vision pioneer from the deep-learning era; its team and tech absorbed into Nebius.

  • Quilt — the AI that auto-answered RFPs and security questionnaires; went dark mid-2026 with no announcement.

  • SinCode AI — the all-in-one AI writing suite; wound down by its own team, users pointed elsewhere.

Want the autopsy? We dug into what actually kills AI tools — funding, acquisition, or just a dead domain — in our data report.

Ethics Corner

When the lab pauses its own model

There's something genuinely new in this week's Astra story, and it's not the hacking — it's the restraint. OpenAI classified its own model as "Critical," then chose not to ship its most powerful cyber capabilities widely, gating them to a vetted few. Anthropic did a version of the same thing, splitting its release into a safe public "Fable" and a locked-down "Mythos" for approved labs. For once, the companies looked at what they'd built and flinched.

That's worth crediting — and worth scrutinizing. Self-classification means the same company that profits from a model also grades its danger and decides who gets the sharp edge. It's the honor system with trillions on the line. And "gated to vetted organizations" only works if the gate holds; this very week showed how fast a determined attacker turns a helpful AI into an accomplice by simply lying to it.

The encouraging read: the offense we saw and the defense (see this week's collection) are advancing together, and the labs are pausing of their own accord. The sober read: a voluntary pause is a decision, not a guarantee — and the moment a rival ships, every pause gets re-litigated. Trust, but verify. Especially when the thing you're verifying can find a zero-day faster than you can.

Quick Tips and Tricks

Prompt of the week: the 10-minute security tune-up

In a week where AI found zero-days and criminals hijacked an agent, the boring basics protect you more than anything fancy. Paste this in and knock it out:

Act as my personal security coach. Walk me through a 10-minute hardening checklist for a regular person, in priority order: (1) which of my accounts most need passkeys or app-based 2FA right now and how to turn them on; (2) how to check if my email/passwords have been in a breach and what to do if they have; (3) the 3 riskiest habits most people have and the quickest fix for each; (4) one thing to set up today that would most limit the damage if a device or account is compromised. Keep each step concrete and skimmable.
AI Events Calendar

On the calendar:

  • Ray Summit — early November, San Francisco. The event for teams building and scaling AI infrastructure.

  • NeurIPS 2026 — December, San Diego. The research world's flagship machine-learning conference.

  • AI for Good Global Summit — the UN ITU's summit on pointing AI at global problems; sessions run through the year.

Until next time 👋

Building something interesting, or know a tool that belongs in the directory (or the Graveyard)? Hit reply, or email [email protected] — we read everything. (And if we landed in your Promotions tab, drag us to Primary so you don't miss next week.)

— The ToolDirectory.AI team