Field notes

Blog — Page 52

Engineering notes, case studies, and lessons from the field.

OpenAI's GPT-5.4 Adds Native Computer Use — Scores 75% on OSWorld, Beating Average Humans
AI Tools

OpenAI's GPT-5.4 Adds Native Computer Use — Scores 75% on OSWorld, Beating Average Humans

GPT-5.4 brings native computer-use capabilities to OpenAI's API and Codex, letting agents control a mouse and keyboard across applications. It scored 75% on OSWorld-Verified, above the average human benchmark of 72.4%.

Mar 31, 20265 min read
Meta Open-Sources TRIBE v2: A Brain Encoding Model That Predicts fMRI Responses at 70x Scale
AI Models

Meta Open-Sources TRIBE v2: A Brain Encoding Model That Predicts fMRI Responses at 70x Scale

Meta's FAIR team released TRIBE v2, a trimodal foundation model trained on 451 hours of fMRI data from 720 subjects. It predicts neural activity across 70,000 individual brain voxels — a 70x jump over the original — and supports zero-shot generalization to new subjects.

Mar 31, 20265 min read
Google Releases Gemini 3.1 Flash Live: Real-Time Voice AI for Agents in 90+ Languages
AI Models

Google Releases Gemini 3.1 Flash Live: Real-Time Voice AI for Agents in 90+ Languages

Google's Gemini 3.1 Flash Live collapses the traditional transcribe-reason-synthesize pipeline into a single native audio-to-audio model. It scored 90.8% on ComplexFuncBench Audio and is now available to developers via the Gemini API.

Mar 31, 20265 min read
Anthropic Confirms 'Mythos' — Its Most Powerful Model Yet — After Accidental Leak
AI Models

Anthropic Confirms 'Mythos' — Its Most Powerful Model Yet — After Accidental Leak

Anthropic's next-generation model, internally codenamed Mythos or Capybara, was exposed through a misconfigured content management system. The company confirmed it represents a 'step change' in capabilities — and that its cybersecurity potential is unprecedented.

Mar 31, 20265 min read
Amazon's Health AI Is Now Open to Every U.S. Customer — No Prime, No One Medical Required
AI Tools

Amazon's Health AI Is Now Open to Every U.S. Customer — No Prime, No One Medical Required

Amazon has fully opened its Health AI assistant to all U.S. customers after a phased rollout that began in January. Built on Amazon Bedrock and One Medical's clinical infrastructure, the agent can interpret lab results, manage prescriptions, book appointments, and connect users to a doctor via direct message — for free with Prime.

Mar 30, 20265 min read
ARC-AGI-3 Offers $2M to Any AI That Can Think Like a Human — Every Frontier Model Scores Below 1%
AI Models

ARC-AGI-3 Offers $2M to Any AI That Can Think Like a Human — Every Frontier Model Scores Below 1%

The ARC Prize Foundation just released its hardest benchmark yet: an interactive reasoning test where GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro all score under 1%, while untrained humans score 100%. The $2M prize remains unclaimed.

Mar 30, 20265 min read
Google Launches Gemini 3 Deep Think: 41% on Humanity's Last Exam, Live for Ultra Subscribers
AI Models

Google Launches Gemini 3 Deep Think: 41% on Humanity's Last Exam, Live for Ultra Subscribers

Google's Gemini 3 Deep Think is now available to AI Ultra subscribers, hitting 41.0% on Humanity's Last Exam and 45.1% on ARC-AGI-2 — using parallel reasoning to explore multiple hypotheses simultaneously.

Mar 30, 20265 min read
Mistral AI Takes On €830M in Debt to Build Sovereign AI Data Center Near Paris
AI Infrastructure

Mistral AI Takes On €830M in Debt to Build Sovereign AI Data Center Near Paris

Mistral AI secured €830 million in debt financing — its first ever — to construct a large-scale AI data center near Paris targeting ~14,000 Nvidia GPUs. The move is a direct bet on European AI infrastructure independence.

Mar 30, 20265 min read
Mistral Small 4 Is One Model That Does the Job of Three — and It's Fully Open Source
AI Models

Mistral Small 4 Is One Model That Does the Job of Three — and It's Fully Open Source

Mistral's new 119B Mixture-of-Experts model unifies reasoning, multimodal understanding, and agentic coding under a single Apache 2.0 license. With only 6B active parameters per token and a 40% latency reduction in optimized configs, it's the most capable open-weight model Mistral has shipped.

Mar 30, 20265 min read

Page 52 of 61 · 549 articles