The Daily Brief · Issue 8 · 25 August 2026Nvidia puts its Groq 3 LPX inference chip into full production as a mystery coding model racks up 26 trillion tokens
Hot Chips week brought Nvidia's fastest inference hardware yet and a push to run CUDA on RISC-V. Anthropic hired the founder of Google's TPU programme for its own chip effort, and the anonymous Ox Alpha model broke usage records on OpenCode.
ChipsNvidia's Groq 3 LPX inference accelerator enters full production
Nvidia announced at Hot Chips 2026 that its Groq 3 LPX chip, built on technology licensed from Groq, is in full production. LPX racks sit beside Vera Rubin systems: Rubin GPUs handle reading long prompts, while LPX speeds up decoding, the step that writes out each token. Nvidia showed Vera Rubin NVL72 with LPX generating 3,400 tokens per second on Gemma 4 31B with a 100,000-token context, measured by Artificial Analysis and described as the fastest ever on that model. Nvidia claims a 4x faster response time than the nearest alternative platform.
Why it matters: Nvidia is answering Cerebras and Etched directly on speed, which should make fast agent responses cheaper to offer.
Source: Wccftech
Model watchAnonymous Ox Alpha processes 26 trillion tokens in four days on OpenCode
OpenCode, an open-source terminal coding agent, said users ran 26 trillion tokens through the stealth model Ox Alpha in its first four days, across 327,000 users and 8.3 million sessions. That made it OpenCode's second most-used model, behind DeepSeek V4 Flash at 33 trillion, and RuntimeWire reports it broke an OpenRouter launch record. Ox Alpha is free for a limited time, has a 1-million-token context window, and its maker is still listed as 'Unknown'. RuntimeWire notes 93% of input tokens came from cache, so the headline number overstates fresh work.
Why it matters: A free, capable model can win developers in days without any brand, showing how weak loyalty to model makers has become.
Source: RuntimeWire
PeopleAnthropic hires Google TPU founder Amir Salek to build its own chips
Anthropic hired Amir Salek, who founded Google's custom-chip programme and delivered its first seven TPU generations, RuntimeWire reports, citing Bloomberg. He joins Anthropic's compute team under James Bradbury. Salek earlier led Nvidia's system-on-a-chip design group and has spent four years as an investor at Cerberus. Anthropic confirmed in August it was recruiting a custom-silicon team. It still trains and runs Claude on Nvidia GPUs, Google TPUs and Amazon Trainium, with large TPU commitments from 2027, so any in-house chip is years away.
Why it matters: Owning its chips could eventually cut Anthropic's cost of serving Claude and reduce its dependence on Nvidia, Google and Amazon.
Source: RuntimeWire
Legal AIThomson Reuters explains how it built Thomson, its own legal AI model
Thomson Reuters, owner of Westlaw and the Reuters newswire, published how it built the Thomson LLM. The model starts from a leading open-weight base model and is further trained on the company's legal, tax and news content with help from its subject experts. The team came from Safe Sign Technologies, a pre-revenue start-up it bought in August 2024. Thomson Reuters says the model is competitive with frontier models at a fraction of their size and cost, and scores 0.914 on instruction following, ahead of Claude Opus 4.8, Gemini 3.1 Pro and GPT-5.5 in its own published tests.
Why it matters: Big data owners are now training their own models on open bases rather than renting general ones, which could change legal research tools.
Source: Thomson Reuters
ChipsNvidia sets out what RISC-V chips need to run CUDA
At Hot Chips, Nvidia described its plan to bring CUDA, its dominant GPU programming platform, to RISC-V CPUs. RISC-V is an open chip design standard that anyone can use without licence fees. CUDA currently supports only x86 and Arm processors. Nvidia's requirements include server-grade RISC-V features, ACPI support for hardware discovery and power management, and cache-coherent PCIe links so data moved to and from the GPU stays accurate. Chips and Cheese notes most existing RISC-V hardware will not meet the bar yet.
Why it matters: Opening CUDA to RISC-V gives chip makers a cheaper CPU option for AI servers while keeping them tied to Nvidia GPUs.
Source: Chips and Cheese
FundingKeenable raises $26 million to build a web search index for AI agents
Keenable, founded by former Yandex search chief Andrey Styskin and AI scientist Matthias Petri, came out of stealth with $26 million in seed funding led by Accel. It says it has built a web index of more than 100 billion documents designed for AI systems, which read far more of each page than people do. Its API is already used by several unnamed AI labs and inference providers. The timing follows Google and Microsoft closing their public search APIs. Keenable is also building a Web Query Language to combine facts from many sources.
Why it matters: AI search tools need an independent web index now that Bing's API is gone, and new suppliers could lower costs for smaller players.
Source: TechCrunch