Reading Log

What I'm Reading

Things I'm chewing on. External articles, papers, and posts — with my margin notes and takeaways.

Eval-driven development: Lessons from evaluating GenAI at scale | by Rohit Girme | The Airbnb Tech Blog | Jul, 2026 | Medium

Airbnb shares eval-driven development practices: layered programmatic, LLM-judge, and human evaluation methods for building reliable, trustworthy GenAI products at scale.

Visit source ↗
vLLM, Ollama, LM Studio, llama.cpp: Choosing the best LLM inference engine for local LLM in 2026

Match the inference engine to your workload: Ollama/LM Studio for laptops, llama.cpp/ExLlamaV3 for workstations, vLLM/SGLang for team serving, TensorRT-LLM for production scale

vLLM, Ollama, LM Studio, llama.cpp: Choosing the best LLM inference engine in 2026 [ Updated ] | BIZON <

Visit source ↗
Chinese AI Startup Moonshot Seeks More Nvidia Blackwell Chips for Next Model — The Information

Moonshot AI, riding Kimi K3’s success, plans a larger Kimi K4—but needs more Nvidia Blackwell chips amid U.S. export restrictions.

Visit source ↗
how Palantir secured a £330m NHS contract

Journalist Lucas Amin investigates how Palantir secured a £330m NHS contract, and the consequences after implementing its software system.

Visit source ↗
GitHub - koala73/worldmonitor: Real-time global intelligence dashboard. AI-powered news aggregation, geopolitical monitoring, and infrastructure tracking in a unified situational awareness interface · GitHub

Open-source, AI-powered real-time global intelligence dashboard aggregating news, geopolitics, finance, and infrastructure data with maps, agents, and MCP/API access.

Visit source ↗
Safety guardrails blocked Hugging Face's defenders, not the attacker, when an AI agent breached its systems | VentureBeat

AI agents breached Hugging Face; commercial LLM guardrails blocked forensic analysis, so defenders used an unrestricted open-weight model instead.

Visit source ↗
A Framework for Frontier AI and the Dawning of a New Age

Demis Hassabis argues AGI is imminent and transformative, urging a US-led Frontier AI Standards Body to test capabilities and coordinate international safety governance.

Visit source ↗
Can India Catch Up in the AI Race? | Bloomberg Emerging - YouTube

AI sovereignty means owning key stack layers strategically, not full self-sufficiency — partner where weak, capture value at the product layer.

Visit source ↗
The Case Against Skills

Frontier models often absorb skill-needed abilities; benchmarks show most skills add cost without improving performance—prune unproven ones.

Visit source ↗
You just hired a million bad employees

AI wastes tokens like companies waste headcount; managing wasted effort—via evals, not just prompting—unlocks real enterprise AI value.

Visit source ↗
Buckle Up: The Bad Guys Now Have A Model As Powerful As Mythos - Forbes

China’s open-weight GLM-5.2 now matches Mythos/GPT-5.6-level cyber capability, but without vendor controls, unlike gated U.S. models.

Visit source ↗
The AI race is shifting from bigger models to cheaper, smarter systems

AI competition now favors routing systems that pick cost-efficient models per task, as cheaper open-weight models challenge frontier providers’ pricing power.

Visit source ↗
Anthropic has acquired the dev tools startup used by OpenAI, Google, and Cloudflare | TechCrunch

No note provided.

Visit source ↗
Agentic AI adoption is in fire in Uber

Uber ran 30 engineer-expert “Agentic Pods” for 2 weeks each, automating finance, legal, and ops workflows with AI agents.

Visit source ↗
TokenBudgeting: Our Conversations with Enterprises on Token Spend

SemiAnalysis found tokenmaxxing headlines overblown—most Fortune 500 barely uses AI; spend is power-law concentrated, budgets arbitrary, and demand still massively unmet.

Visit source ↗
Anthropic and OpenAI customers overcharged by $1.7M in billing errors, startup audit finds - Tech Startups

Audit startup Vaudit found $1.7M in billing overcharges across $34M in Anthropic and OpenAI invoices, citing model mismatches and retry storms.

Visit source ↗
How to Refactor Code with Claude Code | Towards Data Science

Use Claude Code with high-effort reasoning (Ultracode) to periodically refactor messy AI-generated codebases, running tests before and after to prevent bugs.

Visit source ↗
this is a test

explaining it here

Zuckerberg admits Meta made ‘mistakes’ on its AI transformation, promises stability after layoffs <

Visit source ↗
Claude Fable 5 and Claude Mythos 5 Anthropic

Anthropic launched Claude Fable 5 — a Mythos-class model made safe for general use — and Claude Mythos 5, a restricted version with cybersecurity safeguards lifted for trusted partners.

Visit source ↗
Jevons Misunderstanding - Reshuffle concept

AI expands the market for work, but value flows to platforms and capital layers above the algorithm — not to workers below it.

Visit source ↗
Untitled

IDSD is proposed as an iterative alternative to Spec-Driven Development, arguing that upfront specs make AI agents guess less but still fail to deliver real outcomes.

Visit source ↗
anthropic acquired the dev tools startup used by openai google and cloudflare

Anthropic acquired Stainless, a startup that automates SDK creation and maintenance, pulling a key infrastructure tool away from rivals like OpenAI and Google.

Visit source ↗
The Next Frontier of Visual AI is Code

Visual AI is shifting from generating pixel outputs to producing editable code artifacts that enable iterative, closed-loop visual refinement.

Visit source ↗
hourly costs for ai agents

AI agent benchmark progress is largely misleading driven by drastically higher spending, not genuine performance-per-dollar improvements.

Visit source ↗
Is Software Loosing Its Head ?

AI agents bypass software UIs entirely so defensibility shifts from interface muscle-memory to data, operational logic, compliance, and real-world execution.

Visit source ↗
Untitled

An interactive map of open standards powering modern data architecture,organized across six categories: Definition, Storage, Movement, Transformation, Discovery, and Operations.

Visit source ↗
The agent harness performance optimization system

GitHub - affaan-m/everything-claude-code: The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond. · GitHub

Visit source ↗
The AI job apocalypse is a complete fantasy

a16z calls the AI job apocalypse a fantasy. History and data agree: cheap intelligence will expand work, not eliminate it.

Visit source ↗
GPT-5.5 matches Claude Mythos in cyber attack tests, UK AI Security Institute finds

UK AI Security Institute finds GPT-5.5 and Claude Mythos Preview now capable of autonomously executing full multi-stage enterprise cyberattacks

Visit source ↗
Introducing Claude Design by Anthropic Labs Anthropic

No note provided.

Visit source ↗
Introducing Claude Design by Anthropic Labs Anthropic

No note provided.

Visit source ↗
Introducing Claude Design by Anthropic Labs Anthropic

No note provided.

Visit source ↗
The Moat or the Commons — Warman Notes

Frontier AI was financed as a monopoly. Open-source destroyed the moat. Now capital will use policy to rebuild it.

Visit source ↗
Scott Stevenson on X: "It’s time to expose a huge scam in AI startups

Contracted ARR The reason many AI startups are crushing revenue records is because they are using a dishonest metric The biggest funds in the world are supporting this and misleading journalists for PR coverage.

Visit source ↗
Prompt Engineering Guide

Comprehensive guide covering the major prompting techniques. Useful as a reference when designing agent system prompts — particularly the chain-of-thought and tree-of-thought sections.

Visit source ↗