Top story
M5 Ultra makes 512GB local AI possible; it does not make it economical https://www.apple.com/newsroom/2026/08/apple-introduces-m6-and-m5-ultra-for-a-big-leap-in-performance-and-ai-compute/
Apple’s M5 Ultra offers up to 512GB of unified memory. Its largest configuration pairs 36 CPU cores and 80 GPU cores with 1.2TB/s of memory bandwidth. Apple says that capacity can hold LLMs with hundreds of billions of parameters entirely on-device. That is the useful delta behind the “AI compute” pitch: some local experiments that previously needed model offload can now fit in one desktop memory pool. The separate M6 tops out at 32GB, despite its dual 16-core Neural Engine and 170GB/s bandwidth.
Capacity is not throughput, model quality, or total cost. In a 657-comment HN thread, SilverSlash asks who genuinely needs the 512GB configuration; SXX notes that moving from 96GB to 256GB already costs $4,000 in the US. I would treat M5 Ultra as an architectural option when data must remain local, not a default buying recommendation. For workloads above 100B parameters, wait for tokens-per-second, power, and 512GB pricing. If a smaller model already passes acceptance tests, more memory is expensive headroom.
Two tabs I kept
Jalapeño wins an early benchmark; production still waits for 2027 https://newsletter.semianalysis.com/p/openai-jalapeno-better-than-nvidia
SemiAnalysis says it watched A0 silicon run InferenceX in OpenAI’s lab, including more than 700 tokens per second per user for DeepSeek R1 at concurrency one. The caveats decide the verdict: OpenAI supplied every number, the full suite and AgentX were not run, and production ramps during 2027. Impressive ASIC evidence, not a reason to redesign infrastructure yet.
Never publish a VM image that has held a token https://cua.ai/docs/how-to-guides/sandbox/minecraft
cua.ai recovered a Mojang JWT, profile name, and UUID from pagefile.sys, log-file slack, and freed clusters after deleting accounts.json and zero-filling free space. The portable rule for agent sandboxes is excellent: build the distributable image before login; do not clone a credentialed machine. HN’s orbital-decay adds a fair warning that specialized tools may beat computer use anyway.
One connection
Local AI is splitting into two purchases: memory to hold a model, or watts to serve its tokens.
— Tin