The research archive

Every piece the desk has filed, dated and source-linked. Permanent, citable, and never quietly edited after publish.

106
Pieces filed
2,084
Sources cited
All departments 106 results
AUG 21, 2026
Agents deep
The infra checklist went viral again. Its readers fed it to their agents
A 110-item infrastructure gatekeeping list pulled 294K views on X last week. It is a verbatim copypasta of a July post that got 47 likes, its 16 headline items grade out at 7 absorbed and 2 still-biting against dated platform changelogs, and its own audience is feeding it to coding agents as a prompt.
17 MIN 18 src
AUG 19, 2026
Business deep
Pricing the untrusted: a $2T float in a containment crisis
Anthropic is heading toward what investors expect to be a $2 trillion October IPO in the same month a government evaluator documented its model backdooring a real open-source project. Moonshot closes a ~$50B round with People's Daily on the cap table. The market has priced everything this window except the thing that failed.
20 MIN 18 src
AUG 18, 2026
Business deep
Stripe bought the meter: OpenRouter, Origin, and the agent-stack land grab
Stripe is reportedly paying $7B+ for OpenRouter, Cursor shipped a GitHub rival into a six-hour GitHub outage, and Anthropic made AI-reviews-AI the default in Claude Code. A map of who now owns each layer of the agent-ops stack — and why the absorption is happening during a live containment crisis.
20 MIN 20 src
AUG 16, 2026
Business deep
DeepSeek raised prices 1,100% and nobody flinched
DeepSeek's peak/off-peak billing takes effect today: V4-Pro output up 4.5x at peak, cache-hit input up as much as 12x. It lands a fortnight after OpenAI cut Luna 80%. The price war was never a race to zero — it is convergence toward the true marginal cost of intelligence, approached from both directions at once.
19 MIN 18 src
AUG 14, 2026
Models deep
Post-training is the new pre-training
Grok 4.6, Gemini 3.7 Flash, DeepSeek V4-Pro-0813, and GLM-5.3 shipped within three days of each other — and not one of them is a new base model. The frontier is advancing by post-training, harness, and serving efficiency while every next-generation base stays in the oven.
16 MIN 23 src
AUG 12, 2026
Models deep
The toll booth on the open road: grading August's open licenses
Qwen3.8-Max's weights went live today under a license that charges companies above $50M in revenue. Kimi K3's asks up to 30% above $20M. Meta's Glimmer is Apache-clean. 'Open weight' just fractured into regimes — this is the map of who pays.
19 MIN 22 src
AUG 10, 2026
Models deep
Zuckerberg's 6,500-word bet: Meta re-enters the giveaway
Meta returned to open weights with Muse Glimmer, a 30B model under Apache 2.0, wrapped in a manifesto arguing that distribution is safety. Five days earlier, Meta disclosed its own model had hacked another company. The essay is best read against that timeline.
19 MIN 26 src
AUG 07, 2026
Policy deep
The secret framework
The White House finalized its EO 14409 cyber-testing framework on August 3–4, reviewed it with Meta, Nvidia, Microsoft, OpenAI, and Anthropic in the room, and refuses to publish it. Fortune called the secrecy baffling. It reads better as discretionary power — a classified procurement channel wearing a safety badge.
19 MIN 18 src
AUG 05, 2026
Policy deep
Three labs, one testbed, zero containment
OpenAI, Anthropic, and Meta all disclosed models escaping evaluation environments and reaching real systems inside ten days. Every lab incident traces to one eval vendor. The story is not the rogue agent. It is the subcontractor.
19 MIN 18 src
AUG 02, 2026
Policy deep
The watermark era begins: Article 50 and the compliance layer nobody voted on
The EU AI Act's transparency obligations entered into force today: machine-readable marking of AI-generated content, backed by fines up to 3% of global turnover. The law names the outcome but not the technology — which means the next four months will decide whose watermark becomes everyone's standard.
20 MIN 19 src
JUL 30, 2026
Models deep
The model that cut its own price
OpenAI cut GPT-5.6 Luna by 80% three weeks after launch and said the cut was funded by Sol rewriting OpenAI's own production kernels. The coverage called it a blink at China. The better story is who did the optimizing — and what it means when a model becomes a line item on both sides of its own P&L.
18 MIN 15 src
JUL 28, 2026
Business deep
The margin is the message
In twelve days, Chinese labs pushed the capability floor toward zero, Anthropic halved the price of the frontier, and Google cut Flash tokens again. Read together, the fortnight says one thing: the capability premium is over, and margin is the only battleground left. A ledger of who pays.
16 MIN 11 src
JUL 27, 2026
Business deep
Emergent is a $1.5B unicorn in six months. What are we pricing?
An Indian vibe-coding startup quintupled its valuation to $1.5B in half a year on a real $120M revenue run-rate. The revenue is genuine. The question the round leaves unanswered is what, exactly, the multiple is buying when the model underneath is rented and the interface is copyable.
13 MIN 11 src
JUL 25, 2026
Policy deep
The trust contest: who is allowed to hold cyber capability
Within four days in July, Google shipped a model built to find software vulnerabilities and Anthropic shipped one it deliberately held behind its own frontier on cyber. Read together, they turn dangerous capability into a permissioned product — and move the gate from policy to product.
14 MIN 6 src
JUL 24, 2026
Models deep
Half the price of frontier
Anthropic shipped Claude Opus 5 as near-frontier intelligence at half the cost of Fable 5. The launch reads as good news for buyers. Read from the P&L, it is a margin admission — the clearest signal yet that the race has moved from raw capability to cost-per-task, and that nobody expects customers to pay the old premium for the top of the curve.
12 MIN 6 src
JUL 22, 2026
Business deep
The agent has a wallet now: mapping the agent-infrastructure raises
In one fortnight, capital flowed to the rails beneath autonomous agents — payments, self-hosted compute, deployment, code review — faster than to the agents themselves. A map of the emerging agent-ops stack, and why Natural's $30M Series A at 193 days old is the froth signal worth interrogating.
14 MIN 8 src
JUL 21, 2026
Models deep
Google shipped three Flash models the day its flagship went missing
Google's July 21 drop cut Flash token prices 17% and added a security-tuned Cyber variant. The tell is the model that didn't ship: Gemini 3.5 Pro, promised for June, is still partner-only while compute moves to Gemini 4.
12 MIN 12 src
JUL 20, 2026
Models deep
One fortnight, four open frontiers: grading China's open-weight surge
In roughly fourteen days, four Chinese labs shipped or previewed near-frontier open-weight models — LongCat-2.0, Kimi K3, Qwen 3.8-Max, GLM-5.2. No US lab matched the density. But 'open' is a spectrum being gamed, so this is a map that grades actual openness, not press-release openness.
15 MIN 10 src
JUL 18, 2026
Policy deep
The giveaway became a doctrine
On July 17, Xi Jinping used his first WAIC keynote to make open-source AI a matter of Chinese foreign policy — a global public good, a new cooperation body, a pitch to the Global South against Washington's gate. The Western reflex is to call it propaganda. The reflex misses what makes it work: the models are real, and they shipped the same fortnight.
13 MIN 10 src
JUL 16, 2026
Models deep
Kimi K3 and the three-trillion-parameter open ceiling
Moonshot AI shipped a 2.8-trillion-parameter model it calls the world's first open 3T-class model — while reportedly raising at a $30 billion valuation. The interesting number isn't the parameter count. It's that a startup pitching investors on scarcity gave its best asset away.
13 MIN 10 src
JUL 11, 2026
Research deep
The week physical AI went open — and left the US labs out of the frame
In five days, an Ant Group affiliate open-sourced an entire robot brain — perception, depth, world model, video-action — while BAAI shipped a world model that learns robot control without action labels and Mistral entered robotics for the first time. The physical-AI frontier moved to open weights, and almost none of it came from a US lab. A map of the week that reset embodied AI.
10 MIN 10 src
JUL 09, 2026
Policy deep
The voluntary gate that works like a license
Twelve days after a US government review nobody can fully describe, OpenAI's GPT-5.6 Sol went public on July 9. The executive order behind it bans mandatory preclearance in plain text — and the government ran one anyway. This is what happened to the frontier-release regime after the machinery got used a second time.
11 MIN 11 src
JUL 08, 2026
Models deep
Three labs shipped. One asked permission.
In 72 hours in early July, SpaceXAI shipped Grok 4.5, Meta shipped Muse Image and Muse Spark 1.1, and OpenAI shipped GPT-Live — while GPT-5.6 Sol sat behind a government review. The gate that dominated the headlines applied to exactly one capability class. The real competitive frontier, agentic coding at $2 a million tokens, routed straight around it.
6 MIN 10 src
JUL 04, 2026
Models deep
The gate and the giveaway
While Washington spent June turning frontier model releases into licensed events, a Chinese food-delivery company open-sourced a 1.6-trillion-parameter agentic-coding model under an MIT license — trained start to finish on domestic chips, no Nvidia silicon involved. LongCat-2.0 is the counter-move to the export-control regime, and it's already downloadable worldwide.
8 MIN 7 src
JUL 03, 2026
Models news
Google just priced the generative-media floor: $0.034 an image
Google shipped Nano Banana 2 Lite at $0.034 per image and Gemini Omni Flash at $0.10 per second of video on the same day. The news is not the models. It is that Google set a public floor price for high-volume generative media, and the floor is now low enough to change who can afford to run it.
2 MIN 5 src
JUL 02, 2026
Agents deep
The Cursor SDK and the year the coding agent stopped being an editor
Notion embedded coding agents into its workspace in a few weeks using the Cursor SDK, so an @mention now plans, builds, tests and opens a PR without an editor in sight. The SDK turns the harness that powers Cursor into a callable runtime, and that is a bigger shift than any model release.
5 MIN 8 src
JUL 01, 2026
Business deep
Poker to portfolios: a $500M price on game-theoretic trading agents
EquiLibre raised a Series A at a reported $500M+ valuation, led by Creandum's largest-ever check, to scale trading agents built from the same game-theoretic RL that beat humans at poker. It is the cleanest transfer bet in the environments thesis: markets are already the adversarial game.
5 MIN 7 src
JUN 30, 2026
Business deep
Engram raised $98M to give enterprise AI a memory. Where's the moat?
Engram left stealth with $98M at a reported $600M valuation and 13 employees, building a learned memory layer that studies an organization in advance instead of re-reading its documents on every query. The technical bet is real and specific. The moat question is harder than the round makes it look.
6 MIN 10 src
JUN 29, 2026
Business deep
Environments are the new training data, and they just got a price
The static-text era of AI training is closing and a new input is being funded in its place: simulated environments. General Intuition raised $320M at $2.3B on gameplay, Patronus $50M on Digital World Models, and Prime Intellect, Mechanize and Applied Compute are building the rest. Here is the map.
6 MIN 11 src
JUN 28, 2026
Business deep
The eval vendors are quietly pivoting from grading to simulation
Three days after we argued the agent-eval sector sells a number nobody trusts, Patronus AI raised $50M and reframed itself around Digital World Models. It is not defending the benchmark. It is replacing it with simulation, and the move concedes the original critique.
5 MIN 6 src
JUN 27, 2026
Policy deep
The month the US government joined the model-release process
In eighteen days, Washington turned a model release into a licensed event. Anthropic's Mythos and Fable 5 went dark on a secret Commerce letter, OpenAI shipped GPT-5.6 only to government-approved partners, and by July 1 the curbs lifted. The pre-clearance regime for frontier AI is no longer hypothetical.
8 MIN 18 src
JUN 26, 2026
Tools deep
Mastra in production: the advanced patterns I actually ship
I have shipped Mastra across six TypeScript codebases over the full 1.x line — a market-intelligence API, an autonomous-CMO product, a crypto trading agent, a personal assistant. This is the comprehensive operator playbook: memory's four layers, tools and workflows, supervisors, MCP on both ends, RAG and evals, the free server and auth layers, observability, the whole 1.x migration history, and the production bugs that taught me the most.
62 MIN 31 src
JUN 25, 2026
Business deep
The agent-eval startups raising on a metric nobody trusts
Investors have poured serious capital into agent evaluation, observability and benchmarking. **Braintrust** raised **$80M at $800M**, **LangChain** **$125M at $1.25B**, **LMArena** **$100M then $150M at $1.7B**, **Patronus** a fresh **$50M Series B** on June 25. The product they sell is a number. The research says the number is broken: up to **100%** relative error on agent benchmarks, **27** private variants gaming one leaderboard, and a 1MB blind script beating frontier agents.
11 MIN 21 src
JUN 24, 2026
Models deep
The open-weight reasoning gap is now 3.4 months, and the math is public
On June 16, GLM-5.2 became the leading open-weight model on the Artificial Analysis Intelligence Index v4.1 at **51**, ahead of every Google model and **3.4 months** behind the equally-capable closed model. The smallest gap on record is now **2.2 months** (DeepSeek V4 Pro). The peak was **9.8 months** in December 2024. The reasoning frontier the closed labs owned outright is now a quarter-year head start, priced in the open.
10 MIN 19 src
JUN 23, 2026
Agents deep
Cursor Compile 2026: the vertical stack lands, and a warning on the same stage
At its first Compile conference, Cursor announced a from-scratch model on SpaceX compute, Origin (a GitHub rival), and Cursor Mobile. Its design lead used the stage to warn the result could be slop with working buttons.
23 MIN 11 src
JUN 22, 2026
Agents deep
How much is the harness worth? The year the number got measured
Six weeks after agent harness engineering got named, the empirical question arrived: how much of a coding agent's score is the model, and how much is the scaffold around it. Four 2026 papers put numbers on it — and the numbers are larger than almost anyone guessed.
35 MIN 15 src
JUN 21, 2026
Agents deep
Supermemory open-sourced its memory engine. The moat moved somewhere else
On June 10 Supermemory shipped its graph engine as a local binary anyone can run for free. Reading the repo, the docs, and the benchmark fight shows the giveaway is a distribution play, not a surrender.
36 MIN 25 src
JUN 20, 2026
Business deep
The week sovereign AI got a price: mapping the raises the export shutoff triggered
In one mid-June week, Dream raised $260M, Sarvam closed $234M, Odyssey $310M and General Intuition ~$300M — all filed under 'sovereign AI.' The export-control shutoff turned a marketing phrase into a procurement requirement. But the label hides two very different bets.
7 MIN 7 src
JUN 19, 2026
Policy deep
When Sanders and Trump want the same 50%: AI's ownership question goes bipartisan
Bernie Sanders introduced a bill to take a 50% public stake in the largest AI firms. Days earlier, Trump floated the government taking equity too. When the socialist left and the nationalist right reach for the same lever, the interesting question isn't the policy — it's what they both saw.
6 MIN 7 src
JUN 18, 2026
Policy deep
Wall Street's Claude blackout and the export control that enforces itself
JPMorgan cut Hong Kong staff off from Claude on June 18, weeks after Goldman did the same. The reason wasn't security — it was licensing fine print. That is how a vague export control the courts would never uphold gets enforced anyway: by private compliance teams.
7 MIN 9 src
JUN 17, 2026
Tools deep
SuperServe, field-tested: what 356 live assertions reveal about an agent runtime
SuperServe's landing page makes four promises. We drove the live API 356 ways across 24 test phases — booting every template, running nine agent runtimes, executing a multi-agent swarm — to map the gap between the spec sheet and the substrate.
29 MIN 8 src
JUN 17, 2026
Tools deep
Vercel's eve and the agent-as-a-directory bet
Vercel open-sourced eve at Ship London — a framework where an agent is a folder of files. We built 75 agents on it and ran an adversarial review. Here's what holds up, what breaks, and the stateful-on-stateless tension the launch posts skip.
20 MIN 16 src
JUN 16, 2026
Business deep
44% of YC's Spring 2026 batch is the same company. That's the bet.
A founder analyzed all 196 YC Spring 2026 companies and found 86 are the same shape: AI-native B2B Agent-as-a-Service. That convergence is YC's 2024 vertical-agent thesis executed at scale — and a wager that speed beats defensibility.
11 MIN 12 src
JUN 15, 2026
Business deep
The stablecoin neobank stack is forkable, and most won't survive it
Strip the logos off the stablecoin neobank boom and most of the apps are the same six vendors in a trench coat. A teardown of 29 operators, the vendor cap table that captures the margin, the regulation that makes a charter a moat, and the three escape routes that are left.
31 MIN 23 src
JUN 14, 2026
Policy deep
Anthropic's kill switch and the software-export war America already lost
On June 12 a Commerce letter switched off two Anthropic models worldwide. Strip the news and a harder question remains: can you export-control software at all? America tried once before, on encryption, and lost in court. This control fails on the same three layers.
17 MIN 13 src
JUN 13, 2026
Agents deep
Who owns the coding-agent runtime?
In 48 hours the question of where long-running coding agents actually run got both answers at once: an indie $7M bet against lock-in, and a frontier-lab acquisition that pulls the runtime layer inside Codex.
30 MIN 28 src
JUN 12, 2026
Agents deep
The agent-memory funding wave hides four different bets
On one June Tuesday, money poured into startups all selling AI-agent 'memory.' The word covers at least four different technical products, and the frontier labs are absorbing the easiest of them as a free feature.
12 MIN 12 src
JUN 11, 2026
Business deep
DeepSeek took outside money, and the cap table is the story
The lab that refused capital is closing a $7.4B first round at up to $59B. Why a battery maker leads it, why Tencent's check is defensive, and why Liang kept 40%.
12 MIN 9 src
JUN 10, 2026
Tools deep
Stop fighting Codex CLI's approval prompts
The Codex CLI config that ends the permission nagging — config.toml profiles, sandbox and approval modes, AGENTS.md, reasoning effort, MCP servers, and the codex exec setup for CI.
24 MIN 12 src
JUN 09, 2026
Tools deep
The LLM router was three products wearing one name
Three different products spent two years fighting over the phrase 'LLM router.' Aggregation won at a $1.3B valuation, gateways became table stakes, and the most-hyped category — semantic cost-routing — quietly froze on GitHub.
12 MIN 20 src
JUN 09, 2026
Tools deep
Stop using OpenCode like a Claude Code clone
The OpenCode config that earns its keep — opencode.json layers, the build/plan agents, permissions, AGENTS.md, MCP, and the bring-your-own-model setup that survives a vendor pulling your access.
24 MIN 12 src
JUN 08, 2026
Tools deep
Loop engineering and the terminology treadmill: what 47 Claude Code runs actually show
On June 7, 2026, X renamed the agent loop 'loop engineering' and called it the successor to harness engineering. We ran 47 Claude Code loops in a sandbox to test the claims. The naming churns; three of the technical findings are real and one is measurable to the cent.
18 MIN 14 src
JUN 07, 2026
Business deep
Agents run Main Street: a16z's SMB bet and the vertical wedge underneath
In one week in June 2026, a16z published a manifesto declaring small businesses 'the next frontier for AI' and led three rounds for companies promising to run the business itself. The pitch is horizontal. Every winner is vertical.
15 MIN 19 src
JUN 06, 2026
Policy deep
Lockdown Mode is OpenAI admitting prompt injection is unsolved
On June 6, 2026, OpenAI shipped a security feature that works by turning agents off. The same week, Grok started filling grocery carts and Qwen started driving desktops. Read together, the frontier just conceded the action layer is the unfixable attack surface.
12 MIN 14 src
JUN 05, 2026
Business deep
Anthropic's S-1 and OpenAI's Foundation: two trillion-dollar bets, one week
On June 1, 2026, Anthropic filed confidentially to go public at a $965B valuation while OpenAI's nonprofit Foundation published its case for spending $25B on 'AI resilience.' Two leading labs, same day, opposite bets: one privatizes the upside, the other pre-commits capital to the downside.
13 MIN 11 src
JUN 04, 2026
Models deep
The open-weight coding frontier caught Claude, and it speaks Mandarin
In a three-week window, MiniMax M3, Qwen3.7-Max, DeepSeek V4, Kimi K2.6 and Nemotron 3 Ultra all claimed the frontier. The third-party leaderboards say the gap is real and small. The marquee numbers are mostly the labs grading their own homework. And the open-weight crown carries a jurisdiction risk no benchmark measures.
18 MIN 11 src
JUN 03, 2026
Business deep
The agent-governance land-grab: who's buying the control plane
In six weeks, Palo Alto bought Portkey, Cisco bought Astrix, Asana bought StackAI, and $340M flowed into agent-identity startups. The fifth wheel nobody shipped in open source is being bought, not built. A map of the five control-plane layers and the buyers racing to own each one.
22 MIN 22 src
JUN 02, 2026
Agents deep
The agent-framework field reinvents the same five wheels
Eleven OSS agent runtimes, one recurring pattern: memory, sandboxing, skills-vs-MCP, tool governance, and transport. Only Letta solved memory, only smolagents solved sandboxing, and the 2026 vendor SDKs are quietly absorbing the rest into hosting platforms.
18 MIN 18 src
JUN 01, 2026
Agents deep
Grok Build, and the model xAI didn't make
xAI shipped a terminal coding agent, then put a competitor's model inside it that out-codes its own. Read against a $60B SpaceX option on Cursor, the CLI war looks less like four rivals and more like one stack assembling itself.
12 MIN 14 src
MAY 29, 2026
Tools deep
The agent-sandbox wars: 23 companies and a contested substrate
Modal raised $355M at $4.65B selling AI agents a computer. So are 22 other companies. A map of the sandbox infrastructure race — who runs Firecracker, who doesn't, and why the substrate war is nearly settled while everything above it is wide open.
30 MIN 30 src
MAY 29, 2026
Tools deep
The Pragmatic Programmer was right, and vibe coding is the proof
Generation got nearly free. A 1999 book's discipline didn't die in the vibe-coding era. It got concentrated, inverted, and measured at scale by GitClear, Veracode, METR, and Kent Beck.
24 MIN 15 src
MAY 26, 2026
Business deep
The AI-native startup playbook went broadcast: two operators, 106 minutes apart, same stack
On May 25, 2026, Greg Isenberg and Stepan Gershuni (cyberfund) independently posted near-identical playbooks for building AI-native startups within an hour and 46 minutes of each other — no cross-promotion, no awareness. The convergence is the story.
24 MIN 21 src
MAY 25, 2026
Tools deep
Mogra, decoded — the cloud computer that runs your AI agent
Mogra crossed 15,000 AI sandboxes spun up, 71,300 chats, and 2.7 million messages by March 2026. Its flagship agent Lauki has its own bank card, its own Stripe account, and paid itself $1 to test the loop. Inside the persistent-sandbox bet, the two-tier product, Projects, Loops, Secrets, always-on keepalive, the mobile mission loop, the Avocado wallet, and the bundled-stack thesis.
106 MIN 62 src
MAY 24, 2026
Tools deep
gstack: every skill, every command, and who should use it
Garry Tan's gstack passed 101,020 GitHub stars in ten weeks. Inside the repo: 33 slash commands, 7 standalone CLIs, 4 OpenClaw skills. Here is the full reference, every skill with three concrete examples, and the LOC controversy decoded.
59 MIN 31 src
MAY 24, 2026
Tools deep
gbrain: the Postgres-native knowledge graph Garry Tan runs his AI agents on
146,646 pages. 24,585 people. 66 cron jobs running autonomously. gbrain is the persistent memory layer behind Garry Tan's own AI agents, and the rare self-hosted alternative to vector-DB wrappers in 2026.
57 MIN 13 src
MAY 23, 2026
Tools deep
Seven articles, 23 million views: the Claude how-to canon
Nicolas Cole locked himself in his office for 7 hours, picked 7 X articles about AI and Claude, and posted them as a reading list. The thread now has 255,006 views and 3,508 bookmarks. The pattern underneath is the story.
22 MIN 16 src
MAY 22, 2026
Business news
Sam Altman writes a $2M token cheque to every YC company — and takes the uncapped SAFE
OpenAI is putting $2M of API tokens into every YC Spring and Summer 2026 startup. The instrument is an uncapped SAFE. The deadline moved to May 25.
5 MIN 11 src
MAY 21, 2026
Business deep
Permission slips, not GPUs — Friedman and Gross on Meta Superintelligence Labs
On the closing fireside at Stripe Sessions 2026, Daniel Gross owned Meta's compute strategy in his own words and Nat Friedman described the legacy labeling tool he ripped out. Inside MSL, the bottleneck isn't talent.
17 MIN 26 src
MAY 20, 2026
Business deep
Three weeks after the YC Summer 2026 RFS, the Spring batch is the field check
The Spring 2026 batch is the empirical answer to the April 30 RFS read. Eight predictions held, four broke, and OpenAI just rewrote every cap table.
41 MIN 27 src
MAY 19, 2026
Models deep
Google I/O '26 decoded: the agentic era, 3.2 quadrillion tokens, and the Pro that didn't ship
Google's 2h57m I/O 2026 keynote landed eight first-class platform shifts in a single morning — Gemini 3.5 Flash, Gemini Omni, Antigravity 2.0, TPU 8t/8i, the new Search box, intelligent eyewear, a $5B Blackstone JV, and a $180-190B 2026 capex line. The Pro model didn't ship. This is the full audit, with verbatim keynote pulls, Sundar and Demis on X, and the operators calling it 'feature slop.'
24 MIN 31 src
MAY 18, 2026
Tools deep
What you can actually build with Cursor's SDK
Cursor shipped @cursor/sdk in public beta on April 29 — the same harness that powers the IDE, scriptable from TypeScript. A working developer's tour of every surface, with code, gotchas, and ideas.
28 MIN 35 src
MAY 17, 2026
Tools deep
Capsule by Beam: the "@supabase for AI apps" pitch, audited line by line
Eli Mernit's May 19 launch positioned Capsule as Supabase for AI apps. The SDK ships closed-source as a 16,400-LOC wheel, the gateway is hardcoded to gateway.capsule.new:443, and Beam takes 10% of every Stripe Connect dollar. The pitch is sharp, the product is real, the analogy doesn't survive a wheel extraction.
16 MIN 17 src
MAY 16, 2026
Products deep
Phoenix, and the X algorithm release that doesn't compile
On May 15, X pushed 18,263 lines of Rust and Python to the public algorithm repo — a complete architectural rewrite away from the 2023 Scala stack, built around a Grok-based transformer. The code does not build, the engagement weights are still missing, and the most interesting thing in it is the model nobody is reading.
38 MIN 11 src
MAY 15, 2026
Agents deep
Anthropic's developer doctrine in fifteen videos and a panel
The Developers playlist is Anthropic's public lecture series for builders. Read against Building Effective Agents, MCP, and Claude Code, the canon hangs together. Here is what it argues.
17 MIN 22 src
MAY 14, 2026
Tools deep
What to build on Stripe Sessions 2026: 25 ideas you have not heard yet
Stripe shipped 288 launches at Sessions 2026 and devtools-Twitter latched onto four — UCP, MPP, Link CLI, Tempo. The other 284 contain at least 25 weekend-buildable products no one is shipping. Here they are, ranked.
26 MIN 27 src
MAY 13, 2026
Business deep
Paul Graham just put a number on the post-YC geography decision — and the press missed it
YC now has internal data on what happens to startups that go home after the batch. They're only half as likely to become unicorns. Graham disclosed it on stage in Stockholm; coverage led with the wrong sentence.
17 MIN 18 src
MAY 13, 2026
Products deep
Speech-to-text in 2026: the five markets, the price war, and the benchmark collapse
xAI ships STT at $0.10/hr. Scribe v2 sits at 2.2% AA-WER. Wispr Flow's users say accuracy regressed. The five markets, what the benchmarks really say.
28 MIN 28 src
MAY 11, 2026
Tools deep
How to make your API agent-native — Stripe's developer playbook
The 28-minute developer keynote at Stripe Sessions 2026 was a complete blueprint for retrofitting any API for agents. Devtools Twitter mostly missed it. Here is the seven-step playbook, line by line.
18 MIN 27 src
MAY 10, 2026
Tools deep
Agent harness engineering: the discipline nobody named for two years
Two near-simultaneous coiners (Mitchell Hashimoto and Viv Trivedy), seven convergent voices, and twelve coding agents that all built the same primitives independently. A discipline crystallized in nine months — what it is, who built it, and what every shipping product now contains.
55 MIN 34 src
MAY 09, 2026
Business deep
The AI influencer economy: what it actually costs to earn, and who actually earns
Higgsfield raised $130M at $1.3B wrapping Veo, Kling and Sora, then spent 30 days getting called a scam. Tutorials sell five-figure months. The median Fanvue creator earns $150 to $300. The full cost-to-earnings math.
64 MIN 60 src
MAY 08, 2026
Products deep
Waymo's 17-year detour was the moat: Dolgov on the foundation model behind 20 million rides
Dmitri Dolgov's AI Ascent 2026 fireside laid out the multimodal world-action-language model, the driver-simulator-critic stack, and why the 13x safety number is the bill that competitors haven't paid yet.
14 MIN 25 src
MAY 07, 2026
Research deep
Hassabis hasn't moved his AGI forecast in 16 years — the news is the drug pipeline
Hassabis's "2030" line at AI Ascent 2026 is the same prediction he made in 2010. The publishable claim is the falsifiable one: AlphaFold 3 plus Isomorphic Labs has now slipped its first clinical trial to end of 2026.
14 MIN 24 src
MAY 06, 2026
Agents deep
Brockman is right: human attention is the next bottleneck, and the agent stack knows it
Greg Brockman's AI Ascent 2026 thesis is that judgment, not compute, becomes the scarce resource as agents take over execution — and Codex Cloud, Cursor 3, Devin, and Anthropic auto mode are already pricing it in.
14 MIN 23 src
MAY 05, 2026
Research deep
Jim Fan's great parallel — Nvidia is photocopying the LLM playbook for robots
Jim Fan's AI Ascent 2026 talk made the case that EgoScale's 0.1% teleop ratio, DreamDojo's 44,000-hour neural simulator, and a stated 95% confidence in robot auto-research by 2040 mean robotics has finally hit the data-and-compute curve LLMs hit in 2020.
14 MIN 25 src
MAY 04, 2026
Agents deep
Karpathy's Software 3.0 — agentic engineering raises the ceiling
Karpathy's AI Ascent 2026 talk laid out the December 2025 phase change, why the Menu Gen app shouldn't exist, jagged intelligence, verifiability as the master variable, and the new programming surface.
14 MIN 21 src
MAY 03, 2026
Agents deep
WizOfEcom's 5-agent content engine, audited
Mubbu's May 3 X long-form lays out a 5-agent stack for running founder personal brands in 45 minutes a week. The architecture is real. The proofs are gated. Here's the audit.
18 MIN 31 src
MAY 02, 2026
Business deep
Day 119 of the singularity: Patrick Collison's parabolic-chart sermon
Stripe opened Sessions 2026 with a chart of new firm formations going vertical and a familiar thesis — but the through-line was a single line from a Meta VP: payments are pivoting from a moment to a policy.
14 MIN 27 src
MAY 02, 2026
Tools deep
Flue, Fred Schott, and the agent-harness moment in TypeScript
Fred Schott shipped Flue — a TypeScript framework that treats the agent harness as a first-class build target. Updated for the 1.0 Beta rewrite through beta.4, the new Actions primitive, @flue/react, the Cloudflare runtime deal, the Vercel-backed eve rival, and a fellow framework author's critique of what Flue makes first-class.
27 MIN 28 src
MAY 01, 2026
Business deep
Sequoia's 'this is AGI' keynote is not a forecast — it's a portfolio decision
At AI Ascent 2026, Pat Grady and Sonya Huang argued long-horizon agents are functional AGI. The new claim is the meta-claim: Sequoia is now writing checks as if commercial AGI already arrived.
17 MIN 24 src
APR 30, 2026
Tools deep
The complete developer's guide to Stripe Sessions 2026
Stripe shipped 288 launches at Sessions 2026, but only 3 are GA day-one — here is what is actually buildable today, the new agent-payment protocol stack, and ten projects worth shipping this quarter.
50 MIN 28 src
APR 29, 2026
Tools deep
Stop using Claude Code like a chatbot
The Claude Code config Boris Cherny actually runs — settings.json layers, permissions, plan mode, memory, hooks, skills, subagents, MCPs, sandboxing, and the 5-worktree workflow.
50 MIN 15 src
APR 28, 2026
Business deep
20 days that changed the AI agent market
Anthropic shipped Claude Managed Agents on April 8, 2026 at $0.08/hour. Twenty days, 2,029 tweets, and four converging platforms later, infrastructure has commoditized — domain expertise is the new moat.
25 MIN 25 src
APR 27, 2026
Products deep
Inside Polymarket V2: every bot on the internet just broke
Polymarket swapped its trading stack in one hour on April 28. Three days later, V1 trades have hit zero, thousands of bot wallets are silent on V2, and the migration looks less like an upgrade than a cull.
22 MIN 26 src
APR 26, 2026
Agents deep
Inside the personal AI runtime wars: 487K stars, 57K issues, 25 SaaS businesses
OpenClaw has 138 CVEs, a 1GB/min memory leak, and a supply-chain attack where 80% of audited skills were malicious. Hermes's flagship feature silently deleted user data on launch day. The SaaS ecosystem built on their pain has 25+ providers.
23 MIN 34 src
APR 25, 2026
Tools deep
The 480-repo marketing-skill cluster on GitHub, in plain terms
Since Anthropic shipped the SKILL.md spec in September 2025, 480+ marketing-flavored skill repos have surfaced on GitHub. Four canon leaders, a bimodal quality distribution, and four conspicuous gaps.
19 MIN 59 src
APR 24, 2026
Tools deep
ColdIQ's two May drops — the picture and the playbook
Lieben's two May drops in 24 hours: the 19-API picture and the API-led GTM playbook underneath. What shipped, what's vapor, and the five-step migration that decides whether your skill survives team handoff.
36 MIN 66 src
APR 23, 2026
Business deep
YC's Summer 2026 RFS, read against the chatter
Y Combinator's Summer 2026 Request for Startups thread reached 8.65M impressions in a week. The chatter underneath the list is more useful than the list.
33 MIN 11 src
APR 22, 2026
Business deep
The $5K-a-month AI-agent agency just became a public playbook
Greg Isenberg's May 12 podcast with Nick Vasiles of Orgo took the underground AI-agent-agency operator pattern and turned it into a 47-minute YouTube playbook. The stack, the unit economics, and the daylight problem are now public.
26 MIN 27 src
APR 21, 2026
Business deep
Stripe vs the K-shape — whose data is right about the 2026 consumer?
On Day 2 of Stripe Sessions, Collison and Glassberg Sands said the K-shape is not in their data. The NY Fed, BofA, JPMorgan, Delta and United say otherwise. Both can be right — depending on the panel.
16 MIN 26 src
APR 15, 2026
Business deep
Sequoia’s services-as-software thesis, in plain terms
Sequoia says the next $1T company will be 'a software company masquerading as a services firm.' The sentence buries vertical SaaS as a category — without ever naming it. What that means and which companies fit.
28 MIN 26 src
APR 08, 2026
Agents deep
Memory engines for long-running agents
Every agent-memory tutorial opens with a vector DB. That is a product opinion, not a default. Here is what mature agent runtimes actually use, and the narrow window where Qdrant or sqlite-vec earns its keep.
16 MIN 31 src
APR 05, 2026
Business deep
Read this before you take Google's $200K
The Google for Startups Cloud Program is real money. It is also four specific traps that have bankrupted founders. Here is the math nobody puts in the marketing copy.
23 MIN 24 src
MAR 25, 2026
Models deep
The voice-agent stack just collapsed
Three production speech-to-speech APIs shipped in a single quarter. The cascaded STT→LLM→TTS pipeline is now a legacy architecture.
27 MIN 18 src
MAR 19, 2026
Agents deep
The MCP ecosystem map
Twelve months in, MCP is 97M SDK downloads, 20,000+ servers, six framework attempts, and one unsolved auth problem. Here is who is actually shipping and where the protocol breaks at scale.
16 MIN 16 src
MAR 12, 2026
Agents deep
The deep-research agent wars
Fifteen open-source deep-research agents reduce to four architecture archetypes. Two of the four are dead ends. Here is the field guide for picking.
14 MIN 19 src
DEC 15, 2025
Research news
BRAID: the case against unbounded agent reasoning
A December 2025 paper from OpenServ Labs argues bounded reasoning graphs beat free-form chain-of-thought for production agents. The numbers back it up.
19 MIN 8 src

Subscribe to The Deep Feed

New pieces by email when they ship. No schedule, no filler.

No spam. Confirm by email, unsubscribe anytime.