<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="4.4.1">Jekyll</generator><link href="https://blog.palcu.net/feed.xml" rel="self" type="application/atom+xml" /><link href="https://blog.palcu.net/" rel="alternate" type="text/html" /><updated>2026-08-01T16:30:44+01:00</updated><id>https://blog.palcu.net/feed.xml</id><title type="html">Alex Palcuie’s Blog</title><subtitle>Alex Palcuie&apos;s blog. Notes from keeping Claude reliable at Anthropic: model launches, incidents, forecasting, and the odd rant. Written from London.</subtitle><entry><title type="html">Fable 5</title><link href="https://blog.palcu.net/2026/07/fable-5.html" rel="alternate" type="text/html" title="Fable 5" /><published>2026-07-01T23:00:00+01:00</published><updated>2026-07-01T23:00:00+01:00</updated><id>https://blog.palcu.net/2026/07/fable-5</id><content type="html" xml:base="https://blog.palcu.net/2026/07/fable-5.html"><![CDATA[<p>When Anthropic launched Fable 5, the first Mythos-class model to reach general availability, on June 9th, I <a href="https://x.com/AlexPalcuie/status/2064394659510751299">tweeted</a>:</p>

<blockquote>
  <p>launching claude fable 5, a mythos-class model, to general availability is a landmark event in human history that deserves sombre consideration</p>
</blockquote>

<p>Three days later, on June 12th, the US government issued an export control directive and <a href="https://www.anthropic.com/news/fable-mythos-access">Fable 5 access was suspended</a>.</p>

<p><a href="https://www.anthropic.com/news/redeploying-fable-5">Eighteen days later</a>, on June 30th, after a new safety classifier tested with the Department of Commerce blocked the reported technique in over 99% of cases, the export controls were lifted. Fable 5 came back today.</p>

<p>I tweeted <a href="https://x.com/AlexPalcuie/status/2072402841373749455">“we are SO BACK”</a>, and then, once the rate limits were reset, <a href="https://x.com/AlexPalcuie/status/2072429602866184643">“summer holiday is over!”</a>.</p>]]></content><author><name></name></author><summary type="html"><![CDATA[When Anthropic launched Fable 5, the first Mythos-class model to reach general availability, on June 9th, I tweeted: launching claude fable 5, a mythos-class model, to general availability is a landmark event in human history that deserves sombre consideration Three days later, on June 12th, the US government issued an export control directive and Fable 5 access was suspended. Eighteen days later, on June 30th, after a new safety classifier tested with the Department of Commerce blocked the reported technique in over 99% of cases, the export controls were lifted. Fable 5 came back today. I tweeted “we are SO BACK”, and then, once the rate limits were reset, “summer holiday is over!”.]]></summary></entry><entry><title type="html">Claude Sonnet 5</title><link href="https://blog.palcu.net/2026/06/claude-sonnet-5.html" rel="alternate" type="text/html" title="Claude Sonnet 5" /><published>2026-06-30T12:00:00+01:00</published><updated>2026-06-30T12:00:00+01:00</updated><id>https://blog.palcu.net/2026/06/claude-sonnet-5</id><content type="html" xml:base="https://blog.palcu.net/2026/06/claude-sonnet-5.html"><![CDATA[<p>From the <a href="https://www.anthropic.com/news/claude-sonnet-5">announcement</a>:</p>

<blockquote>
  <p>Claude Sonnet 5 is built to be the most agentic Sonnet model yet. It can make plans, use tools like browsers and terminals, and run autonomously at a level that, just a few months ago, required larger and more expensive models.</p>
</blockquote>

<p>Close to Opus 4.8 performance at Sonnet prices.</p>]]></content><author><name></name></author><summary type="html"><![CDATA[From the announcement: Claude Sonnet 5 is built to be the most agentic Sonnet model yet. It can make plans, use tools like browsers and terminals, and run autonomously at a level that, just a few months ago, required larger and more expensive models. Close to Opus 4.8 performance at Sonnet prices.]]></summary></entry><entry><title type="html">Anthropic Series H</title><link href="https://blog.palcu.net/2026/05/anthropic-series-h.html" rel="alternate" type="text/html" title="Anthropic Series H" /><published>2026-05-28T18:00:00+01:00</published><updated>2026-05-28T18:00:00+01:00</updated><id>https://blog.palcu.net/2026/05/anthropic-series-h</id><content type="html" xml:base="https://blog.palcu.net/2026/05/anthropic-series-h.html"><![CDATA[<p>From the <a href="https://www.anthropic.com/news/series-h">announcement</a>:</p>

<blockquote>
  <p>Anthropic has raised $65 billion in Series H funding led by Altimeter Capital, Dragoneer, Greenoaks, and Sequoia Capital, valuing the company at $965 billion post-money.</p>
</blockquote>

<p>Run-rate revenue crossed $47 billion in May. Three and a half months after <a href="/2026/02/anthropic-series-g.html">Series G</a>, the chart needs a new y-axis.</p>]]></content><author><name></name></author><summary type="html"><![CDATA[From the announcement: Anthropic has raised $65 billion in Series H funding led by Altimeter Capital, Dragoneer, Greenoaks, and Sequoia Capital, valuing the company at $965 billion post-money. Run-rate revenue crossed $47 billion in May. Three and a half months after Series G, the chart needs a new y-axis.]]></summary></entry><entry><title type="html">Claude Opus 4.8</title><link href="https://blog.palcu.net/2026/05/claude-opus-48.html" rel="alternate" type="text/html" title="Claude Opus 4.8" /><published>2026-05-28T12:00:00+01:00</published><updated>2026-05-28T12:00:00+01:00</updated><id>https://blog.palcu.net/2026/05/claude-opus-48</id><content type="html" xml:base="https://blog.palcu.net/2026/05/claude-opus-48.html"><![CDATA[<p>From the <a href="https://www.anthropic.com/news/claude-opus-4-8">announcement</a>:</p>

<blockquote>
  <p>We’re upgrading Claude Opus to a new version: Claude Opus 4.8. It builds on Opus 4.7 with improvements across benchmarks, and is a more effective collaborator. It’s available today for the same price.</p>
</blockquote>

<p>Same price, and fast mode is now three times cheaper than it was for previous models. I <a href="https://x.com/AlexPalcuie/status/2060078625546731799">tweeted</a>:</p>

<blockquote>
  <p>brought to you by our cracked inference engineers</p>
</blockquote>]]></content><author><name></name></author><summary type="html"><![CDATA[From the announcement: We’re upgrading Claude Opus to a new version: Claude Opus 4.8. It builds on Opus 4.7 with improvements across benchmarks, and is a more effective collaborator. It’s available today for the same price. Same price, and fast mode is now three times cheaper than it was for previous models. I tweeted: brought to you by our cracked inference engineers]]></summary></entry><entry><title type="html">SpaceX / Anthropic compute deal</title><link href="https://blog.palcu.net/2026/05/spacex-anthropic-compute-deal.html" rel="alternate" type="text/html" title="SpaceX / Anthropic compute deal" /><published>2026-05-20T12:00:00+01:00</published><updated>2026-05-20T12:00:00+01:00</updated><id>https://blog.palcu.net/2026/05/spacex-anthropic-compute-deal</id><content type="html" xml:base="https://blog.palcu.net/2026/05/spacex-anthropic-compute-deal.html"><![CDATA[<p>Anthropic <a href="https://www.anthropic.com/news/higher-limits-spacex">announced</a> a compute partnership with SpaceX earlier this month:</p>

<blockquote>
  <p>We’ve agreed to a partnership with SpaceX that will substantially increase our compute capacity. This, along with our other recent compute deals, means that we’ve been able to increase our usage limits for Claude Code and the Claude API.</p>
</blockquote>

<p>Colossus 1 in Memphis adds more than 300 megawatts, over 220,000 NVIDIA GPUs, within the month. Bloomberg <a href="https://www.bloomberg.com/news/articles/2026-05-20/anthropic-to-pay-spacex-nearly-45-billion-for-computing-deal">reports</a> the deal has since expanded to Colossus 2, nearly $45 billion through May 2029. Our cofounder Tom Brown <a href="https://x.com/nottombrown/status/2057194829986300375">tweeted</a>:</p>

<blockquote>
  <p>We’re expanding our partnership with @SpaceX, and will be scaling up on GB200 capacity in Colossus 2 throughout June.</p>

  <p>Appreciate @elonmusk and the team helping us find good homes for the Claudes.</p>
</blockquote>

<p>There is also a mention of multiple gigawatts of orbital compute.</p>]]></content><author><name></name></author><summary type="html"><![CDATA[Anthropic announced a compute partnership with SpaceX earlier this month: We’ve agreed to a partnership with SpaceX that will substantially increase our compute capacity. This, along with our other recent compute deals, means that we’ve been able to increase our usage limits for Claude Code and the Claude API. Colossus 1 in Memphis adds more than 300 megawatts, over 220,000 NVIDIA GPUs, within the month. Bloomberg reports the deal has since expanded to Colossus 2, nearly $45 billion through May 2029. Our cofounder Tom Brown tweeted: We’re expanding our partnership with @SpaceX, and will be scaling up on GB200 capacity in Colossus 2 throughout June. Appreciate @elonmusk and the team helping us find good homes for the Claudes. There is also a mention of multiple gigawatts of orbital compute.]]></summary></entry><entry><title type="html">Amazon / Anthropic compute deal</title><link href="https://blog.palcu.net/2026/04/amazon-anthropic-compute-deal.html" rel="alternate" type="text/html" title="Amazon / Anthropic compute deal" /><published>2026-04-20T12:00:00+01:00</published><updated>2026-04-20T12:00:00+01:00</updated><id>https://blog.palcu.net/2026/04/amazon-anthropic-compute-deal</id><content type="html" xml:base="https://blog.palcu.net/2026/04/amazon-anthropic-compute-deal.html"><![CDATA[<p>Anthropic <a href="https://www.anthropic.com/news/anthropic-amazon-compute">announces</a> an expanded partnership with Amazon:</p>

<blockquote>
  <p>We have signed a new agreement with Amazon that will deepen our existing partnership and secure up to 5 gigawatts (GW) of capacity for training and deploying Claude, including new Trainium2 capacity coming online in the first half of this year and nearly 1GW total of Trainium2 and Trainium3 capacity coming online by the end of 2026.</p>
</blockquote>

<p>Over one million Trainium2 chips already train and serve Claude, and Anthropic is committing more than $100 billion to AWS over the next ten years.</p>]]></content><author><name></name></author><summary type="html"><![CDATA[Anthropic announces an expanded partnership with Amazon: We have signed a new agreement with Amazon that will deepen our existing partnership and secure up to 5 gigawatts (GW) of capacity for training and deploying Claude, including new Trainium2 capacity coming online in the first half of this year and nearly 1GW total of Trainium2 and Trainium3 capacity coming online by the end of 2026. Over one million Trainium2 chips already train and serve Claude, and Anthropic is committing more than $100 billion to AWS over the next ten years.]]></summary></entry><entry><title type="html">Claude Mythos Preview</title><link href="https://blog.palcu.net/2026/04/claude-mythos-preview.html" rel="alternate" type="text/html" title="Claude Mythos Preview" /><published>2026-04-07T12:00:00+01:00</published><updated>2026-04-07T12:00:00+01:00</updated><id>https://blog.palcu.net/2026/04/claude-mythos-preview</id><content type="html" xml:base="https://blog.palcu.net/2026/04/claude-mythos-preview.html"><![CDATA[<p>Today we <a href="https://www.anthropic.com/glasswing">announced</a> Claude Mythos Preview as part of Project Glasswing. It scores 77.8% on SWE-bench Pro, up from 53.4% for Opus 4.6.</p>

<p>My reliability team was asked for feedback on it, so from page 204 of the <a href="https://www-cdn.anthropic.com/08ab9158070959f88f296514c21b7facce6f52bc.pdf">model card</a>:</p>

<blockquote>
  <p>From a reliability engineering perspective, the model still cannot be left alone in a production environment to use generic mitigations. It frequently mistakes correlation with causation and it is not able to course-correct for different hypotheses. When asked to write incident retrospectives, more often than not it focuses on a single root cause and does not consider multiple contributing factors. However, we’ve found this model to be a step change in two areas. The first is signal gathering and initial analysis, where, by the time an engineer has opened two dashboards, the model has already found the outliers and what’s breaking. The second case is navigating ambiguity when there is a clearly defined outcome. For example, due to time zone differences, the reliability team in London was asked to stand up a model in a production environment with different constraints, and the engineers were unfamiliar with both the task and the constraints. Claude Mythos Preview was able to work step-by-step, fixing each error by observing other environments, checking any breadcrumbs that were left in previous commits, and reading documentation.</p>
</blockquote>

<p>The London team in question was us.</p>]]></content><author><name></name></author><summary type="html"><![CDATA[Today we announced Claude Mythos Preview as part of Project Glasswing. It scores 77.8% on SWE-bench Pro, up from 53.4% for Opus 4.6. My reliability team was asked for feedback on it, so from page 204 of the model card: From a reliability engineering perspective, the model still cannot be left alone in a production environment to use generic mitigations. It frequently mistakes correlation with causation and it is not able to course-correct for different hypotheses. When asked to write incident retrospectives, more often than not it focuses on a single root cause and does not consider multiple contributing factors. However, we’ve found this model to be a step change in two areas. The first is signal gathering and initial analysis, where, by the time an engineer has opened two dashboards, the model has already found the outliers and what’s breaking. The second case is navigating ambiguity when there is a clearly defined outcome. For example, due to time zone differences, the reliability team in London was asked to stand up a model in a production environment with different constraints, and the engineers were unfamiliar with both the task and the constraints. Claude Mythos Preview was able to work step-by-step, fixing each error by observing other environments, checking any breadcrumbs that were left in previous commits, and reading documentation. The London team in question was us.]]></summary></entry><entry><title type="html">One year at Anthropic: $2B to $30B run-rate</title><link href="https://blog.palcu.net/2026/04/one-year-at-anthropic.html" rel="alternate" type="text/html" title="One year at Anthropic: $2B to $30B run-rate" /><published>2026-04-06T09:00:00+01:00</published><updated>2026-04-06T09:00:00+01:00</updated><id>https://blog.palcu.net/2026/04/one-year-at-anthropic</id><content type="html" xml:base="https://blog.palcu.net/2026/04/one-year-at-anthropic.html"><![CDATA[<p>From the <a href="https://www.anthropic.com/news/google-broadcom-partnership-compute">Google and Broadcom partnership announcement</a>:</p>

<blockquote>
  <p>Our run-rate revenue has now surpassed $30 billion—up from approximately $9 billion at the end of 2025.</p>
</blockquote>

<p>I joined Anthropic on April 1, 2025. Around that time, CNBC <a href="https://www.cnbc.com/2025/05/16/anthropic-ai-credit-facility.html">reported</a>:</p>

<blockquote>
  <p>Annualized revenue reached $2 billion in the first quarter, the company confirmed, more than doubling from a $1 billion rate in the prior period.</p>
</blockquote>

<p>$2 billion to $30 billion. 15x in the year I have been here.</p>]]></content><author><name></name></author><summary type="html"><![CDATA[From the Google and Broadcom partnership announcement: Our run-rate revenue has now surpassed $30 billion—up from approximately $9 billion at the end of 2025. I joined Anthropic on April 1, 2025. Around that time, CNBC reported: Annualized revenue reached $2 billion in the first quarter, the company confirmed, more than doubling from a $1 billion rate in the prior period. $2 billion to $30 billion. 15x in the year I have been here.]]></summary></entry><entry><title type="html">Claude Code source leak</title><link href="https://blog.palcu.net/2026/03/claude-code-source-leak.html" rel="alternate" type="text/html" title="Claude Code source leak" /><published>2026-03-31T18:00:00+01:00</published><updated>2026-03-31T18:00:00+01:00</updated><id>https://blog.palcu.net/2026/03/claude-code-source-leak</id><content type="html" xml:base="https://blog.palcu.net/2026/03/claude-code-source-leak.html"><![CDATA[<p>VentureBeat <a href="https://venturebeat.com/technology/claude-codes-source-code-appears-to-have-leaked-heres-what-we-know">has the story</a>. The Anthropic statement:</p>

<blockquote>
  <p>Earlier today, a Claude Code release included some internal source code. No sensitive customer data or credentials were involved or exposed. This was a release packaging issue caused by human error, not a security breach. We’re rolling out measures to prevent this from happening again.</p>
</blockquote>

<p>I <a href="https://x.com/AlexPalcuie/status/2039229774422306867">tweeted</a>:</p>

<blockquote>
  <p>I repeat this to every new joiner at Anthropic but it’s worth repeating in public – we have a blameless culture and no single individual is at fault when bespoke complex systems break at scale</p>
</blockquote>

<p>My colleague Jake Eaton sent me the long version of that argument. The NTSB asks why, not who, and that’s <a href="https://asteriskmag.com/issues/05/why-you-ve-never-been-in-a-plane-crash">why you’ve never been in a plane crash</a>.</p>

<p>On a smaller note, the April 1st surprise got spoiled too.</p>]]></content><author><name></name></author><summary type="html"><![CDATA[VentureBeat has the story. The Anthropic statement: Earlier today, a Claude Code release included some internal source code. No sensitive customer data or credentials were involved or exposed. This was a release packaging issue caused by human error, not a security breach. We’re rolling out measures to prevent this from happening again. I tweeted: I repeat this to every new joiner at Anthropic but it’s worth repeating in public – we have a blameless culture and no single individual is at fault when bespoke complex systems break at scale My colleague Jake Eaton sent me the long version of that argument. The NTSB asks why, not who, and that’s why you’ve never been in a plane crash. On a smaller note, the April 1st surprise got spoiled too.]]></summary></entry><entry><title type="html">The Register on my QCon London talk</title><link href="https://blog.palcu.net/2026/03/the-register-qcon.html" rel="alternate" type="text/html" title="The Register on my QCon London talk" /><published>2026-03-20T07:00:00+00:00</published><updated>2026-03-20T07:00:00+00:00</updated><id>https://blog.palcu.net/2026/03/the-register-qcon</id><content type="html" xml:base="https://blog.palcu.net/2026/03/the-register-qcon.html"><![CDATA[<p>The Register <a href="https://www.theregister.com/2026/03/19/anthropic_claude_sre/">wrote up</a> my QCon London talk on using Claude for incident response:</p>

<blockquote>
  <p>When Claude is asked to produce a postmortem report, it delivers “an 80 percent story that’s pretty, it’s readable and convincing,” said Palcuie, but “it’s really bad at root causes.” Claude says “this was the thing, and we all know it is not one thing. It’s not one root cause… It was never the rollout. It was never the code change. It was all the processes in the company that allowed the incident. And Claude doesn’t know the history of your system, especially if your system has been there for ten years.”</p>
</blockquote>

<p>Strange to read your own words quoted back at you by El Reg.</p>]]></content><author><name></name></author><summary type="html"><![CDATA[The Register wrote up my QCon London talk on using Claude for incident response: When Claude is asked to produce a postmortem report, it delivers “an 80 percent story that’s pretty, it’s readable and convincing,” said Palcuie, but “it’s really bad at root causes.” Claude says “this was the thing, and we all know it is not one thing. It’s not one root cause… It was never the rollout. It was never the code change. It was all the processes in the company that allowed the incident. And Claude doesn’t know the history of your system, especially if your system has been there for ten years.” Strange to read your own words quoted back at you by El Reg.]]></summary></entry></feed>