Good morning and welcome back to Not a Llama, the daily podcast where we break down the biggest news in AI models. Today is Thursday, October the eighth, two thousand and twenty six. We have a big show for you. Anthropic just shipped Claude Haiku five point five, their cheapest and fastest small model ever. OpenAI claims a math breakthrough with seven hundred and twenty two papers solving ninety of the top five hundred open problems. And ChatGPT is getting an intelligent user interface rolling out alongside GPT six. Let us dive in. Let us get right into the lead story. Anthropic has just shipped Claude Haiku 5.5, billed as the cheapest and fastest small model the company has ever released. This lands on October the seventh of two thousand twenty six. The API identifier is claude-haiku-5-5, and it arrives roughly a year after Haiku 4.5. According to the OpenRouter listing, this model offers a context length of one million tokens, which is a huge jump from its predecessor. On pricing, input tokens cost ten cents per million up to one hundred thousand tokens, and fifty cents beyond. Output tokens run fifty cents up to one hundred thousand tokens, and two dollars fifty after that. But here is the real story. Simon Willison has flagged what he calls a hidden price increase. The new tokenizer uses about one point twenty five times more tokens than Haiku 4.5, so the headline price does not tell the whole tale. Whether this model finally stops being a token devourer is the question on everyone's mind. The benchmark numbers are striking. On Anthropic's own GDPval test, Haiku 5.5 scores 1620 against 735 for its predecessor. On OSWorld, it reaches 72.4 percent compared to 15.7 percent. Humanity's Last Exam without tools lands at 45.9 percent, versus 10.2 percent before. It is clearly a major upgrade in raw capability, even if the tokenizer costs give reason to pause. The model is available now across AWS, Google Cloud, Microsoft Azure, the Claude Platform, and OpenRouter. Anthropic also halved Sonnet 5.5 cache reads to ten cents per million, and is handing out monthly API credits to subscribers. This is a big launch for the small model space. Let us get straight into the headlines that are dominating the AI world today. First, a massive claim from OpenAI. The company published seven hundred twenty-two math manuscripts from an internal frontier model in a public repository. The model itself stays locked away. According to Latent Space, the collection spans three hundred seventy-two families, drawn from about four thousand research problems, using roughly three hours of ChatGPT Pro thinking compute per result on average. Sam Altman called it a new era of discovery. Some commentators are going further, calling it the most significant moment in mathematics in over a century. Reported results include faster integer multiplication and partial progress on famous conjectures, though an analysis estimates about twenty percent are disproofs. Not everyone is convinced. Will Depue expects some results will not survive scrutiny. Now, over to Mistral Large four, nicknamed Le Chonk. It carries one trillion total parameters with forty-nine billion active, and it is natively multimodal. It is live through the API now, with pricing starting at one dollar thirty-six per million input tokens. Open weights are promised by the end of October. Now, things are getting a little more visual over in the OpenAI camp. According to The Verge, ChatGPT is getting what it is calling an Intelligent UI, and it is rolling out right alongside GPT six. So what does that actually mean? Well, answers are no longer just plain text. They now combine words with diagrams, charts, forms, and tappable buttons. OpenAI says it trained GPT six on exactly when to reach for an interactive visual instead of plain prose, and how to format those pieces well. The examples sound genuinely handy. Picture a diagram of a seven-speed bicycle, with buttons that highlight different parts as you explore it. Or imagine ChatGPT teaching you Mahjong, generating visuals of each tile and dropping them into scrollable categories. You can even ask for inline tools like a retirement savings calculator, a retro game, or a bill splitter. This all ships to ChatGPT Plus, Pro, Business, and Enterprise users starting today. Access expands to Go and free tiers on Thursday, with higher-tier subscribers getting the Sol model and everyone else leaning on the more efficient Luna version. New contenders keep landing in the small model space, and today LiquidAI made its first open Liquid Foundation Model family available. The release dropped on October seventh, two thousand and twenty six, according to the Hugging Face Blog and the Reddit community r slash LocalLLaMA. Two models are live on Hugging Face, LiquidAI d1 three B and LiquidAI d1 omni six hundred million, with demos you can run in the System One Arcade Space. These are decision models, not generative ones. They answer in a single forward pass instead of producing token by token. The d1 three B claims a forty eight point five seven score on the Decision Index, edging out four billion and nine billion parameter models. Its mean score of eighty two point nine across seven datasets beats the Decider four B. The smaller omni model scores seventy eight point four with a quarter of the parameters. Speed is where this really shines. LiquidAI worked with NVIDIA to clock the three B model answering in sixteen milliseconds on a Jetson AGX Thor and under ten milliseconds on an RTX forty nine zero nine. Both models ship their own code and need transformers version five point fourteen or higher. Mistral is back, and this time it is bringing a serious cluster to the party. According to Matthew Berman, the new Mistral Large four, nicknamed Le Chonk, arrived on 3,800 Nvidia Grace Blackwell GPUs, trained from scratch in a European data center. That is a big infrastructure bet, and the real question is whether it can turn into lasting momentum. The model carries one trillion total parameters with forty-nine billion active ones, and it is natively multimodal. You get a half million token context window and an inference speed of one hundred sixteen tokens per second. Pricing is oddly precise at one dollar thirty-six per million input tokens and four dollar eighteen per million output tokens. A public preview is live, though weights are not out until the end of the month, and it is not in the chat app yet. Benchmarks are strong, with a first place finish on Cyber Gym. But the Intelligence Index shows thirty-eight, last of twenty-five. So the capability is real, the infrastructure is bold, and the durability story is just beginning. Here is story six of seven. On the voxel pagoda test, a head-to-head comparison offers a stark reality check. According to a Reddit post on the r/ClaudeAI subreddit dated the seventh of October, two thousand and twenty six, one user reports that Claude Haiku five point five cost twenty four dollars ninety six. The same task ran on GPT six Luna came in at just one dollar ninety six. That makes Haiku five point five roughly twelve times more expensive on the poster's figures. The token usage tells the story. Haiku five point five consumed two hundred sixty eight point eight million model effort tokens with four point four six million tokens out. GPT six Luna used sixty four point five million effort tokens with one point one eight million out. The poster attributes the jump to API prices quintupling after one hundred thousand tokens of context. Haiku five point five reportedly hit that limit first, while Luna hits its threshold at two hundred seventy two thousand context but auto-compacts. Both models were accessed through their respective subscriptions in atomic.chat. The poster even calls Haiku five point five a token goblin. Yet all figures come from a single unverified post. No independent benchmark confirmation appears in the material. So take the twelve dollar fifty multiplier with a grain of salt. Let us pick up where we left off with a pull request that could matter for anyone running models without a giant graphics card. According to a Reddit post on the r LocalLLaMA subreddit, dated October seventh twenty twenty six, user jacek twenty twenty three shared a call from am seventeen an to the ggml-org llama.cpp repository. It is pull request number twenty nine thousand eight hundred eighty seven. The title asks for a GPU cache for MoE experts that stay in host memory. In other words, experts that live on your regular system memory instead of on your graphics card. The poster describes this as a potentially big speedup for MoE models that don not fully fit in VRAM. That framing is exactly what GPU poor local inference has been waiting for. Still, the post invites people to show their speedups, which means no benchmark numbers, no parameter counts, and no release dates are provided here. We have no license information, no specific model names, and no measured performance figures. The big speedup remains an unverified claim. And only a pull request is referenced, not a merged or released feature. So this is a promising direction, not a finished product. Keep an eye on the discussion if you are squeezing more speed out of local models without upgrading your hardware. Today the takeaway is clear. Small models are getting cheaper and faster while big labs keep pushing harder. Anthropic's Haiku five point five leads that charge, and GPT six follows with a smart new interface. Thank you for listening, and a brand new episode arrives tomorrow morning. Until then, this is Not a Llama.