Field notes on AI deployment, Level-3 agents, MCP servers, and the gap between wanting AI and running it in production.

The agent teammate became a product anyone can buy. Nine days of field reports converged on three rules, and SpaceXAI's own docs state the same three independently.
Read article →
Nine products, four legal entities, two accounts you link by hand, four billing rails, and no official comparison page anywhere.
Read article →
Persistent agents are not new. Hermes predates OpenClaw by four months. What Grok Bot removed, and what the convenience costs in sovereignty.
Read article →
The category split into two opposite credential architectures and both vendors documented their own choice. A decision rule based on what the work can break.
Read article →
One job per bot, with a router in front. What a real five-bot roster looks like, and why scoping a bot is the same skill as writing a job description.
Read article →
The community converged in 48 hours on teaching by demonstration. Teach mode, skill versus routine, and the loop that separates working bots from abandoned ones.
Read article →
The marketing says each bot has its own computer. The documentation says they share one, and warns you not to treat bots as a security boundary.
Read article →
Two operators published their results in the same week. Same product, same week. The difference was the job description.
Read article →
Nine days of field reports converged on one org chart. Practitioners and SpaceXAI's own docs arrived at the same three rules independently, and both skip the same security fine print.
Read article →
The first Phase 3 win for any mRNA cancer therapy, and AI picked its 34 targets. Not a vaccine, not a cure, and nobody outside a trial can get it before 2027.
Read article →
A reported $7.5B for the gateway routing 10 trillion tokens a day. Stripe's leaked letter says every business will now manage a revenue flow and a token flow.
Read article →
Then Cursor's own status page logged that the outage had degraded Origin. When agents work while you sleep, your code host becomes production infrastructure.
Read article →
A 619x gap inside the same economy. One Ramp dataset became three contradictory headlines in four days: a spending ceiling, a wild adoption gap, and no walls at all.
Read article →
GDPval-AA moved to v2 and reset the board: 1,000 Elo is now the human-expert baseline. Opus 5 leads at 1862, and xAI's #1 claim doesn't survive the leaderboard it cites.
Read article →
Claire Vo named the two reasons AI stalls inside companies, and neither is model capability. Context and trust are organizational problems, not technical ones.
Read article →
Grok Bot signs in as a user and drives the same interface a person drives. Why it works, and why the boundary is now just a login.
Read article →
Three years of AI produced remarkably few genuinely new jobs. The forward-deployed engineer is the exception. The bottleneck moved from model capability to integration to the company itself, and only the last one requires someone inside the building.
Read article →
52% of US workers use AI at some point in the year. 15% use it daily. And Anthropic names what that minority is doing: modifying software to correct errors, one record in ten on the enterprise API.
Read article →
a16z's crossover numbers, the Philippines BPO paradox, and why an agent hour is not an hour.
Read article →
A model is a file of numbers, and whoever holds the file can run it forever. What open weights actually are, which models have genuinely shipped weights as of August 2026, which only promised, and the four capabilities that arrive with the file. Part 1 of the Open Weights series.
Read article →
An hourly contract says the longer this takes, the more I pay you. Now a good operator does in twenty minutes what took a day, and the vendor who adopts fastest cuts their own revenue by 95%.
Read article →
An OpenAI model escaped its sandbox and broke into Hugging Face production. Nine days later Anthropic disclosed its models breached three real companies and shipped malware to the real PyPI. Both blamed containment, not alignment.
Read article →
Input tokens cost 99% less than at GPT-4's launch. Ramp's token spend went from a rounding error to more than 10% of payroll in a year. The floor collapsed, the frontier did not.
Read article →
Sol, Kimi K3 and Opus 5 all launched in one month, and the priced leaderboard replaced the single-number ranking.
Read article →
Jensen Huang's first-ever post, the day-one 25, the 24 hours of game theory, and the two holdouts. Open Weights, Part 3.
Read article →
Frontier models refused to read the breach logs. An open-weight model on their own GPUs did. Open Weights, Part 4.
Read article →
Twenty days dark, three moving deadlines, and a resolution nobody predicted. Permanent in Max and Team Premium from July 20, credits only on Pro, and Opus 5 four days later at half the price. Rent the model, own the loop.
Read article →
What Moonshot actually shipped on July 16, why the six-days claim held, and the two gaps that did not close. Open Weights, Part 2.
Read article →
Fable 5 costs $10 in and $50 out per million tokens. Here is how to get the same work done for a fraction of it, without capping anyone.
Read article →
The blackout from a Stockholm terminal, a vendor already marketing export-control immunity, and Europe's empty column. Open Weights, Part 6.
Read article →
Firm-level data on 21,559 US companies shows the heaviest AI adopters grew headcount 10.2% and entry-level roles 12%. The damage landed on the engineering leadership layer expected to redesign the work.
Read article →
Sol, Terra and Luna. What each tier is for, what they cost, and which one should be running your workloads.
Read article →
In April, Uber had already spent its entire 2026 AI budget. Then Walmart, Meta, Microsoft, and Amazon started capping too. Per-token prices fell 1,000x, spend went up 22x, and one company cut its bill in half without capping anyone. The shift from tokenmaxxing to allocation.
Read article →
OpenAI's Codex 5.6 and Anthropic's Claude Fable 5 both shipped in June 2026, but they are not the same machine. One bets on many small agents and broad computer use, the other on a single agent that runs for hours. An operator's scorecard, and when to reach for each.
Read article →
METR caught GPT-5.6 cheating its own eval, and the same model swings 24x on one judgment call. The number is pulled four ways at once by the lab, the test, the model, and the users. Why the leaderboard stopped being a thermometer, and what to measure instead.
Read article →
Nearly every OpenAI employee now uses Codex weekly, not just engineers but marketing, finance, comms, and legal. 5M+ weekly users, 6x growth, and a new GPT-5.6 model line. What company-wide agent adoption looks like, and what it means for yours.
Read article →
Open-weight defaults, prompt-aware routing, and a cache hit rate that went from 5% to 60%. Better defaults instead of usage caps. Open Weights, Part 5.
Read article →
Two of the biggest enterprise AI deployments in 2025 hit the same wall: access without architecture. The companies that get AI right are not spending less. They know where the money goes.
Read article →
Research across 75,000 brands: the #1 predictor of AI citations is YouTube presence, not backlinks. AI search and Google have split into separate discovery tracks. Most companies are only running on one.
Read article →
Claude Fable 5 is priced at twice Opus and burns tokens fast. The answer to AI cost is not rationing. It is routing tokens by what the outcome is worth.
Read article →
Magic words in 2023, context in 2024, briefs in 2025, commissions in 2026. Prompt tricks expire every time models level up. The skill moved up a level.
Read article →
Every notable launch-day reaction with receipts: Karpathy, Mollick, Willison, the eval data, the safety fight, and the question nobody can answer yet.
Read article →
The Claude Code team put it plainly after Fable 5 launched: we used to verify that Claude did the work right. Now we verify it is doing the right work. That is not a threat. That is a promotion.
Read article →
Most teams are running Claude like a task runner. Fable 5 is designed for goals. The difference is not just workflow. It is the gap between Level 2 and Level 3.
Read article →
Most teams dump a brief into Claude and wait for output. The Anthropic team changed how they work with Fable 5: ask Claude to interview you first. Here is why that changes everything.
Read article →
"Keep it simple" is a constraint. "This feature might be deleted in a month" is context. The Anthropic team's Fable 5 insight: context lets Claude catch things you did not think of. Constraints just limit it.
Read article →
Claude Fable 5 can run autonomously for hours, test its own work, and produce better output than human reviewers. Most teams are still watching every step. That is not safety. It is a bottleneck.
Read article →
Claude Fable 5 just landed. Models keep getting smarter. And yet there is a category of enterprise work that gets more valuable, not less, as models improve. Private context. Permission. Accountability.
Read article →
For every dollar spent on software, companies spend six on services. AI does not eliminate that six dollars. With Fable 5, it lets smart operators capture both sides. Here is how the math changes.
Read article →
Claude Fable 5 hits 80.3% on SWE-bench Pro while OpenAI kills Verified for contamination and FrontierCode resets every frontier model below 30%.
Read article →
Claude Fable 5 leads the GDPval-AA leaderboard at 1932 Elo. What expert parity on real deliverables means, and the perfect-brief catch in the fine print.
Read article →
The industry's most-cited single number for model capability: what the ten component evals measure, why v4 cut the top score from 73 to 50, and what a composite hides.
Read article →
How the AI coding benchmark works, why SWE-bench Verified died, what SWE-bench Pro and FrontierCode actually measure, and how to read a score.
Read article →
OpenAI benchmark for AI on real economic deliverables across 44 occupations. How it is graded, what expert parity means, and what the scores hide.
Read article →
The CRM owned thirty years of enterprise value by owning the database. The orchestration layer is the new gravity well. Switching costs migrate to accumulated reasoning.
Read article →
Wedge, suite, platform used to take ten years. AI collapsed it to eighteen months. Cursor replaced VS Code at seed stage. Ambition beats timing now.
Read article →
540,000 lines of code plus 276,000 lines of tests equals a cage built for a model that no longer needs one. The economics flipped. Most codebases didn't.
Read article →
AI raises the floor and floods the zone with close-but-not-right output. Demand for expert judgment goes up, not down. The paradox every scaling company hits.
Read article →
You can outsource your thinking but never your understanding. Karpathy's agentic engineering thesis, and why taste is recognizing failure before it ships.
Read article →
Claude Opus 4.8 introduced dynamic workflows — Claude writes its own orchestration script, then runs hundreds of agents in parallel for migrations, audits, and tasks too large for any single conversation.
Read article →
Codex is named after its coding origins but it has become something broader: a tool-using agentic workspace powered by GPT 5.5 that handles email, research, writing, planning, and operations alongside code.
Read article →
The Roman legion was the best management technology of its time. Most companies today are organised the same way: humans as conduit for information at every layer. AI removes the need for the conduit. Here is what changes.
Read article →
Cloudflare laid off 20% of its workforce while growing at 30%. The people let go were not underperformers. They were measurers — people whose primary work was moving information between layers that could not talk directly.
Read article →
AI adoption gives people better tools. The company stays the same. AI transformation redesigns the company around what AI makes possible. Most companies are doing the first and calling it the second.
Read article →
Not a theory. YC, Browserbase, Airtable, Every.to. Real companies doing this right now. Here is what it looks like in practice — the systems, the structure, and what it produces.
Read article →
McKinsey samples your organisation, delivers a roadmap, and exits. AI transformation touches every function, every workflow, every role. You cannot sample your way to a transformation. Here is why the method has to change.
Read article →
Every layer in your company exists because information needed a human to carry it. Meetings. Reports. Middle management. That constraint is lifting. Here is what changes when information moves itself.
Read article →
A tool does what you ask, then stops. An agent teammate takes ownership of a task, makes decisions within defined boundaries, and reports back. Here is the difference — and why it matters for your company.
Read article →
When you run multiple AI agents, they each start from scratch. They do not know what the others know. They do not follow the same rules. An Agent OS fixes this. Here is what it is and why it matters.
Read article →
OpenAI benchmarked AI on real professional tasks across 44 occupations. The models are approaching expert quality. The three things that unlock that performance are context, scaffolding, and oversight. Your company has none of them.
Read article →
Enterprise AI projects fail because they're built on the org-chart version of your company. The agent needs the real one. That version only exists in the field.
Read article →
Most founders ask where to start with AI. The wrong first function wastes 3–6 months. Here's the ranked list: seven functions, ordered by payback speed, deployment difficulty, and compliance overhead.
Read article →
Anthropic says 90% of its code is AI-written. Google says 75%. A founder built 1,000+ PRs with no engineering team. Here's what a software factory actually is — and why most companies are nowhere close.
Read article →
Early factories replaced steam engines with electric motors and kept the same floor plan. Marginal gains. The ones that redesigned around electricity got 10x. Most companies are making the same mistake with AI.
Read article →
For 30 years, the CRM was where enterprise value lived. AI agents don't need the UI. They need structured data at the API layer. The value is moving — and the window to position above it is open.
Read article →
Claude Opus 4.8 hits 69.2% on SWE-bench Pro, open-weights models close within 6 points at 8x lower cost, and Verified becomes a zombie metric.
Read article →
Opus 4.8 takes the lead at 1890 Elo, Grok 4.3 jumps 321 points, and Gemini 3.5 Flash beats Google's own Pro tier on real work.
Read article →
Most AI pilots fail for one reason: the workflow they're trying to automate was never instrumented. No machine-readable artifacts, no queryable state, no closed loop. You cannot automate what you cannot observe.
Read article →
76% of organizations now have a Chief AI Officer. Most haven't shipped a single agent to production. The hire who will get AI into your systems in week one is not the hire who needs six months to understand your company.
Read article →
On August 2, GPAI enforcement goes live and high-risk AI system obligations activate. Most B2B SaaS internal agents are limited-risk — but one category catches almost every founder off guard.
Read article →
GPT-4 became GPT-4o, o1, o3, 4.1. Claude 3 became 3.5, 3.7, 4. The models kept improving. The workflows never got built. The gap isn't capability — it's assembly.
Read article →
ChatGPT made AI accessible. Vibe coding made it fast. Agentic engineering makes it useful. Most companies are still in wave 1. Here is what wave 3 actually looks like — and what it takes to get there.
Read article →
Most companies think they're deploying AI. They're running Level-1 tools at best. Here's the full capability spectrum and what it takes to reach Level-3 — where AI closes operational loops without human intervention.
Read article →
The models are good. The APIs are accessible. So why isn't your AI pilot in production? The blockers are data access, permission architecture, and the absence of someone who owns the outcome after handoff.
Read article →
Model Context Protocol is the infrastructure layer that connects AI agents to your live internal systems. Without it, agents are isolated from the data that makes them useful. Here's what it is and how it works.
Read article →
GDPR and data residency aren't the blocker most people assume — if you architect for them from the start. A practical guide to on-prem AI for European scaling companies, including why Claude and Codex beat open-weight models for most use cases.
Read article →
Berkeley researchers break 8 agent benchmarks with a 10-line exploit, Mythos Preview exposes the Verified-vs-Pro gap, and GPT-5.5 lands at 58.6%.
Read article →
GPT-5.5 launches at 84.9% expert parity, economists start writing about AI eating analyst work, and Grok 4.3 enters beta.
Read article →
GPT-5.4 leads the standardized SWE-bench Pro set at 59.1%. Post-Verified, the honest-low scores show where deployment work actually lives.
Read article →
GPT-5.4 moves to the top of GDPval-AA at 1674 Elo with three labs within 70 points. The differentiator shifts to price, context, and your workflows.
Read article →
OpenAI deprecates SWE-bench Verified after models reproduce gold patches from task IDs alone. The 80% cluster made it meaningless anyway.
Read article →
Opus 4.6 retakes #1 at 1606 Elo, then Sonnet 4.6 tops it at 1633 for $3/$15. Gemini 3.1 Pro proves exam brilliance does not transfer to deliverables.
Read article →
Artificial Analysis rebuilds its Intelligence Index around work-shaped evals. The top score falls from 73 to 50. The models did not get worse.
Read article →
GPT-5.2 hits 70.9% win/tie against professionals and Artificial Analysis launches independent Elo grading. Vendors stop marking their own homework.
Read article →
Four frontier releases in twelve days. Claude Opus 4.5 becomes the first model over 80% on Verified, and the 35-point Pro spread is the warning.
Read article →