đź“° Good reads

Essays, breakdowns, and explainers worth one focused read. (47 entries)

Linear for agents: tools, evals, harness design

Aishwarya Goel (@aishwarya_08) · Oct 5
Linear for agents - a long thread on tools for AI agents, covering evals, harness design, and production patterns…

Aishwarya Goel's long thread, 'Linear for agents,' covering tools for AI agents: evals, harness design, and production patterns. Linear (the project tool) is invoked as an analogy for the kind of tight, well-designed tooling agents need. A good survey of the agent tooling landscape as of October 2026.

  • Survey of AI agent tools: evals, harness, production
  • 'Linear for agents' sets a design-quality bar
  • Good landscape overview, recent
agentstoolingevals
View original on X ↗

Stop arguing about which AI agent is best

helicerat (@helicerat0x) · Oct 5
stop arguing about which AI agent is best…Claude Cowork…

Helicerat argues people should stop debating which AI agent is best, with images and a Claude Cowork mention. The point: the harness, evals, and workflow matter more than the model badge. Tool debates are usually proxy wars for workflow differences.

  • Agent debates miss the point; harness and workflow matter more
  • Claude Cowork as the current reference
  • Focus on your loop, not the leaderboard
agentsopiniontooling
View original on X ↗

Uber MCP infrastructure

santi (@santtiagom_) · Oct 3
translated from Spanish, Uber MCP infra…

Santi shares a translated post on Uber's MCP infrastructure. MCP (Model Context Protocol) is the emerging standard for connecting agents to tools and data. Seeing how Uber builds MCP infrastructure at scale is a preview of enterprise agent plumbing.

  • Uber-scale MCP infrastructure breakdown
  • MCP is becoming the agent-tool standard
  • Enterprise plumbing preview
mcpuberinfrastructure
View original on X ↗

Bad code and AGENTS.md

Tomas Vykruta (@tvykruta) · Sep 29
long thread on bad code and AGENTS.md…

Tomas Vykruta's long thread on bad code and AGENTS.md. AGENTS.md is the emerging convention of a markdown file that tells AI coding agents how to work in your repo. The thread connects code quality to agent effectiveness: agents amplify whatever codebase you give them, good or bad.

  • AGENTS.md: instructions for AI agents in your repo
  • Agents amplify your codebase's quality, good or bad
  • Clean code matters more, not less, in the agent era
agents-mdcode-qualitycoding-agents
View original on X ↗

System 1 vs System 2 Agent Harnesses

Avi Chawla (@_avichawla) · Sep 28
System 1 vs System 2 Agent Harnesses

Avi Chawla frames agent harnesses as System 1 vs System 2, borrowing the fast/slow thinking distinction. System 1 harnesses react quickly with cached patterns; System 2 harnesses reason deliberately. The framing helps decide when an agent should think fast vs think hard.

  • Fast vs slow thinking applied to harness design
  • Match harness depth to task difficulty
  • Useful framing for your own harness decisions
harnessagentsmental-model
View original on X ↗

DoorDash Vera breakdown

Yarchi (@undefinedKi) · Sep 27
DoorDash Vera breakdown…

Yarchi breaks down DoorDash's Vera, the internal data agent serving 10,000+ employees: curated context, enforced access, real evals, and plan approval. Vera went from 43% to 90% answer accuracy by beating plain Claude Code by 25 points with the same model. You have studied this deeply already; it remains the canonical enterprise agent case study.

  • Vera: 43% to 90% accuracy via context engineering, not a better model
  • The four pillars: curated context, enforced access, real evals, plan approval
  • Canonical case study for enterprise agent architecture
veradoordashenterprise-agentscase-study
View original on X ↗

How to Design an Agent Harness

Yarchi (@undefinedKi) · Aug 15
No text captured (article or link save).

Yarchi's article: six decisions that turn a model into a worker you can leave alone. The harness decisions (tools, memory, evals, recovery, permissions, observability) are the difference between a demo and a system. Six decisions is a manageable framework for designing your own.

  • Six harness decisions for autonomous workers
  • Framework for designing your own harness
  • Directly applicable to your agent runtime work
harnessagentsdesign
View original on X ↗

Xiaomi MiMo V2.6 RL open weights

Adithya S K (@adithya_s_k) · Sep 26
No text captured (article or link save).

Adithya S K shares Xiaomi's MiMo V2.6 RL model on Hugging Face. Open-weight RL-trained models are valuable for research and for builders who want to inspect or fine-tune. Another data point in the open-model wave.

  • MiMo V2.6 RL open weights on Hugging Face
  • Useful for research and fine-tuning
  • Part of the accelerating open-model wave
open-weightsxiaomimodels
View original on X ↗

Ontology is at the center

Chad Wahlquist (@chadwahl) · Sep 24
Ontology is at the center…

Chad Wahlquist argues ontology is at the center, linking Palantir. An ontology (a structured map of what things exist in a domain and how they relate) is what lets AI systems reason about a business instead of just predicting text. This connects directly to your Bloomberg-ontology and Palantir MMDP studies.

  • Ontology as the center of enterprise AI
  • Structured domain maps enable reasoning, not just prediction
  • Connects to your Bloomberg and Palantir research
ontologypalantirenterprise-ai
View original on X ↗

Halo now supports SDPG

Yifan Zhang (@yifanzhang_) · Sep 21
Congrats! Halo now supports…SDPG, arxiv.org/abs/2606.04036

Yifan Zhang announces Halo now supports SDPG, linking a paper (arxiv 2606.04036). SDPG appears to be a training or alignment method; Halo is the framework adopting it. New method support in a popular framework is how research becomes usable.

  • Halo adds support for the SDPG method
  • Framework adoption is how research becomes usable
  • Check the paper for what SDPG actually does
trainingmethodshalo
View original on X ↗

Loop Engineering for Sub-Dime Agents

Google Cloud Tech (@GoogleCloudTech) · Sep 21
No text captured (article or link save).

Google Cloud Tech's article on Loop Engineering for Sub-Dime Agents: making agent loops so cheap each run costs under ten cents. Cost per task is the gating factor for agent deployment; sub-dime loops change what products are viable. Engineering for cost is as important as engineering for capability.

  • Engineering agent loops down to sub-dime cost
  • Cost per task gates real deployment
  • Cheap loops enable new product categories
agentscostengineering
View original on X ↗

SpaceX AI engineer: 2,500 PRs a month

kloess (@kloss_xyz) · Sep 21
SpaceXAI engineer 2,500 PRs/month…

Kloess shares the story of a SpaceX AI engineer shipping 2,500 PRs a month. The number is startling and probably reflects heavy agent assistance. It is a glimpse of what AI-accelerated engineering velocity looks like at the extreme.

  • 2,500 PRs/month: extreme AI-assisted velocity
  • Glimpse of the agent-accelerated future
  • Ask what breaks at that speed, not just what ships
productivityagentsspacex
View original on X ↗

Safety theater

CYBERDOG (@oCYBERDOGo) · Sep 18
long post on safety theater…

CYBERDOG's long post on safety theater: safety measures that look reassuring but do not actually reduce risk. The concept matters because theater consumes the budget and attention that real safety work needs. Read critically; the term gets overused, but the underlying point is real.

  • Safety theater: reassuring measures that do not reduce risk
  • Theater crowds out real safety work
  • Apply the lens skeptically, including to this post
ai-safetycritiqueessay
View original on X ↗

Agent Substrate: K8s infrastructure

Ayaan (@twtayaan) · Sep 17
long post on Agent Substrate K8s infra…

Ayaan writes a long post on Agent Substrate: Kubernetes infrastructure for agents. Kubernetes (K8s) orchestrates containers at scale; an agent substrate is the platform layer agents run on. As agents move to production, the infrastructure questions (scheduling, isolation, scaling) become as important as the model questions.

  • Agent substrate: the platform layer agents run on
  • K8s brings scheduling, isolation, and scaling to agents
  • Infrastructure matters as much as models in production
infrastructurekubernetesagents
View original on X ↗

What OpenAI's agentic software factory looks like

Gergely Orosz (@GergelyOrosz) · Sep 15
Here's what OpenAI's agentic software factory looks like…

Gergely Orosz reveals what OpenAI's agentic software factory looks like. An agentic software factory is OpenAI's internal system for producing software with AI agents at scale. A rare look inside how the frontier lab industrializes agent-assisted coding.

  • Inside OpenAI's agent-powered software production
  • How the frontier lab industrializes coding agents
  • Compare with Uber's software factory piece
openaiagentsengineering
View original on X ↗

RSI algorithms date back to 1987

Juergen Schmidhuber (@SchmidhuberAI) · Sep 13
Concrete algorithms for recursive self-improvement (RSI) date back to 1987…

Juergen Schmidhuber notes concrete algorithms for recursive self-improvement date back to 1987. RSI (systems that improve themselves) is often discussed as a future breakthrough; Schmidhuber points out the formal foundations are decades old. Essential historical context for your RRSI work.

  • RSI formalisms date to 1987, not the future
  • Historical context for self-improvement research
  • Schmidhuber's priority claims are worth knowing
rsihistoryself-improvement
View original on X ↗

Closed-loop and cybernetics

Yi Ma (@YiMaTweets) · Sep 13
long post on closed-loop/cybernetics…

Yi Ma's long post on closed-loop systems and cybernetics, linking to lab materials. Cybernetics is the old science of control and communication in machines and animals; closed-loop means the system's outputs feed back into its inputs. These ideas predate modern AI and quietly underpin RL and agent design.

  • Cybernetics: the science of feedback and control
  • Closed-loop thinking underpins RL and agents
  • Old ideas, newly relevant
cyberneticscontrol-theoryfoundations
View original on X ↗

ToolGrad from Google Research

Google Research (@GoogleResearch) · Sep 10
Introducing ToolGrad…

Google Research introduces ToolGrad. The name suggests gradient-like optimization applied to tool use: improving how agents select and use tools through a learning signal. New agent-training methods from Google Research deserve attention.

  • ToolGrad: optimizing agent tool use
  • From Google Research, so take it seriously
  • Read the linked material for the mechanism
googleagentsresearch
View original on X ↗

AI safety orgs list

Owain Evans (@OwainEvans_UK) · Sep 9
AI safety orgs list…

Owain Evans shares a list of AI safety organizations. The safety ecosystem is fragmented across research labs, nonprofits, and company teams; a map of who is who saves a lot of confusion. Useful context whether or not safety is your focus.

  • Map of the AI safety organization landscape
  • Saves confusion in a fragmented ecosystem
  • Good context even for non-safety work
ai-safetyorganizationsreference
View original on X ↗

Tiny models

Ryan Wright (@IQReactorAI) · Sep 3
long post on tiny-models…

Ryan Wright's long post on tiny models. Small models are getting capable enough for real tasks at a fraction of the cost and latency. The tiny-model trend matters for agents: cheap, fast models change what you can afford to put in a loop.

  • Tiny models are becoming genuinely useful
  • Cost and latency advantages compound in agent loops
  • Watch the capability-per-dollar curve
small-modelsefficiencyagents
View original on X ↗

Stuck in a massive local optimum

Tim Rocktaschel (@_rockt) · Sep 3
We are stuck in a massive local optimum.

Tim Rocktaschel argues the field is stuck in a massive local optimum. A local optimum is a good-but-not-best solution that is hard to leave because every small step looks worse. The provocation: current approaches work well enough to discourage the exploration that could find something better.

  • Are we stuck in a local optimum?
  • Good-enough methods discourage exploration
  • Provocation worth arguing with
opinionresearchprovocation
View original on X ↗

How to RL post-train a 397B model

Edward Hu (@edwardjhu) · Sep 1
How does one RL post-train a 397B model for long-horizon knowledge work?…

Edward Hu asks how one RL post-trains a 397B model for long-horizon knowledge work. Post-training a 397-billion-parameter model with reinforcement learning is at the edge of what is feasible; the infrastructure and algorithmic challenges are extreme. Read for the scale of the problem, not a how-to.

  • RL post-training at 397B parameters
  • Extreme infrastructure and algorithmic challenges
  • Read for perspective on the scaling frontier
post-trainingscalerl
View original on X ↗

Inference-stack rant

Nick Khami (@skeptrune) · Aug 31
long inference-stack rant…

Nick Khami's long inference-stack rant. Rants from practitioners often contain more truth than polished posts: the frustrations reveal where the tooling genuinely hurts. Read for the pain points; they are product and research opportunities.

  • Practitioner frustrations with the inference stack
  • Pain points are product opportunities
  • Rants reveal truth that polished posts hide
inferenceopinionrant
View original on X ↗

Apple Agent Seer paper

elvis (@omarsar0) · Aug 29
long post on Apple Agent Seer paper…

Elvis covers Apple's Agent Seer paper in a long post. Apple publishes sparingly on AI, so their agent research gets extra attention. Worth reading for how a privacy-first, on-device company thinks about agents differently from cloud labs.

  • Apple's take on agent research
  • Privacy-first, on-device perspective differs from cloud labs
  • Rare Apple AI publication; read closely
appleagentspaper
View original on X ↗

Agentic Kernels in Production

Brian Li (@BrianLi23) · Aug 28
No text captured (article or link save).

Brian Li's article on Agentic Kernels in Production. Kernels here likely means the core execution units of agent systems running reliably in production. Production concerns (reliability, observability, cost control) separate real systems from demos.

  • Agentic kernels running in production settings
  • Production concerns: reliability, observability, cost
  • Separates real systems from demos
agentsproductionkernels
View original on X ↗

Running a Software Factory Efficiently at Uber Scale

Uber Engineering (@UberEng) · Aug 28
No text captured (article or link save).

Uber Engineering's article on running a software factory efficiently at Uber scale. A 'software factory' is the industrialized pipeline for producing software: CI/CD, testing, deployment at massive scale. Uber-scale lessons on developer productivity and system reliability.

  • Uber-scale software production practices
  • Industrialized development pipelines
  • Lessons on reliability at scale
engineeringuberscale
View original on X ↗

GLM 5.3 scaling with post-training, explained

Alex Ker (@thealexker) · Aug 28
No text captured (article or link save).

Alex Ker's article explains GLM 5.3's scaling with post-training, intuitively. Post-training scaling (getting more from training after pretraining) is where much current progress comes from. An intuitive explainer beats the paper for first contact.

  • Intuitive explainer of GLM 5.3's post-training
  • Post-training scaling is a key progress lever
  • Read before the paper
glmpost-trainingexplainer
View original on X ↗

Google wiki skills paper

DAIR.AI (@dair_ai) · Aug 28
long post on Google wiki skills paper…

DAIR.AI covers Google's wiki skills paper. Wiki skills likely refers to agents learning skills from wiki-style knowledge. Google's agent research output is prolific; this one is worth a skim for the skill-acquisition angle.

  • Google paper on wiki-derived skills
  • Skill acquisition is a key agent frontier
  • Skim for the core idea
googleskillspaper
View original on X ↗

Meta/Stanford paper breakdown

marfin (@marfinxx) · Aug 26
long post on Meta/Stanford paper…

Marfin's long post breaking down a Meta/Stanford paper. Joint industry-academic papers often combine scale with rigor. The breakdown format saves reading the full paper if you only need the key results.

  • Meta/Stanford joint paper explained
  • Industry scale plus academic rigor
  • Read the breakdown, then decide on the paper
papermetastanford
View original on X ↗

Cursor on reliable Git hosting

Cursor (@cursor_ai) · Aug 18
We're making Git hosting more reliable, performant, and scalable…

Cursor writes about making Git hosting more reliable, performant, and scalable. Version control infrastructure is invisible until it breaks; Cursor's scale makes their lessons valuable. Relevant if you run infrastructure, skimmable otherwise.

  • Reliable Git hosting at Cursor's scale
  • Invisible infra until it breaks
  • Skimmable unless you run infra
gitinfrastructurecursor
View original on X ↗

Thoughts about scaling laws

jietang (@jietang) · Aug 19
long post on Thoughts About Scaling Law…

Jietang's long post sharing thoughts about scaling laws. Scaling laws (predictable relationships between compute, data, model size, and performance) underpin billion-dollar training decisions. Thoughtful commentary helps separate what is established from what is hype.

  • Commentary on scaling law science
  • Separates established results from hype
  • Useful framing for training discussions
scaling-lawstheoryessay
View original on X ↗

DeepSeek harness talk summary

Phosphen (@phosphenq) · Aug 16
long DeepSeek harness talk summary…

Phosphen summarizes a DeepSeek harness talk on video. DeepSeek's engineering culture produces unusually frank technical talks. Harness talks from strong labs are high-signal for your interests.

  • DeepSeek harness talk summarized
  • Strong labs give frank technical talks
  • High signal for harness design
deepseekharnesstalk
View original on X ↗

Irreducible error essay

helicerat (@helicerat0x) · Aug 14
long irreducible error essay…

Helicerat's long essay on irreducible error. Irreducible error is the portion of mistakes no model can eliminate, the noise floor of a task. Understanding it calibrates expectations: some failures are not fixable with better models, only with better task design.

  • Irreducible error: the mistakes no model can fix
  • Calibrates expectations for agent reliability
  • Some failures need task redesign, not better models
errortheoryessay
View original on X ↗

Update on our post-training effort

Gabe Pereyra (@gabepereyra) · Aug 20
No text captured (article or link save).

Gabe Pereyra's article with an update on a post-training effort. Post-training (instruction tuning, RLHF, and friends) is where raw models become useful products. Practitioner updates reveal what actually worked, which is rarer than theory.

  • Practitioner update on post-training work
  • What actually worked, not just theory
  • Compare methods against your own needs
post-trainingpractitionerupdate
View original on X ↗

Microsoft AI Frontiers internship project

Xidulu (@xidulu) · Aug 11
1/ Sharing a new, interesting project I did during my internship at Microsoft AI Frontiers…

Xidulu shares a project from a Microsoft AI Frontiers internship, with a paper link. Internship projects at frontier labs are often surprisingly substantial. Worth a look at what the next generation of researchers is working on.

  • Microsoft AI Frontiers internship project
  • Frontier-lab intern work is often substantial
  • Window into emerging research directions
researchmicrosoftinternship
View original on X ↗

Long essay on Canada

Eliot Pence (@EliotPence) · Aug 6
long essay on Canada…

Eliot Pence's long essay on Canada. Topic is outside your usual AI bookmarks, which suggests it caught your eye for personal reasons. Long essays reward slow reading; skim the opening to decide if it earns the full read.

  • Long-form essay on Canada
  • Outside your usual topics; read for interest
  • Skim the opening to judge the full read
essaycanada
View original on X ↗

5 papers to read this year

ajay yadav (@bettersayAJ) · Aug 2
if you only read 5 more papers this year, make it these: 1. Language Models are Few-Shot Learners (GPT-3) 2. CLIP 3. Retrieval-Augmented Generation…

Ajay Yadav says if you only read 5 more papers this year, make it these: the GPT-3 paper (few-shot learning), CLIP (connecting images and text), RAG (retrieval-augmented generation), and two more. These are foundation papers; reading them is rite-of-passage material for the field.

  • 5 foundation papers: GPT-3, CLIP, RAG, and more
  • Rite-of-passage reading for the field
  • Read the papers, not just summaries
papersfoundationsreading-list
View original on X ↗

Verifiability

Yarchi (@undefinedKi) · Aug 1
long post on Verifiability…

Yarchi's long post on verifiability. Verifiability means being able to check an agent's work independently, not just trust its outputs. For enterprise agents, verifiable outputs are the difference between a demo and something you can deploy with accountability.

  • Verifiable agent outputs enable real accountability
  • Trust-but-verify beats blind trust in production
  • Core requirement for enterprise deployment
verifiabilityagentstrust
View original on X ↗

Harness engineering as a discipline

beamnxw (@beamnxw) · Jul 30
This paper is f*cking brilliant A computer science paper establishes harness engineering as the proper engineering discipline for agents…

Beamnxw shares a computer science paper establishing harness engineering as a proper engineering discipline for agents. The harness (everything around the model: tools, prompts, evals, loops) is where production agent quality actually lives. A paper arguing it deserves discipline-status validates your whole build direction.

  • Paper argues harness engineering is its own discipline
  • The harness, not the model, is where quality lives
  • Academic validation of your build direction
harnessagentspaper
View original on X ↗

Every serious agent will run on a graph

sunick (@Serantych) · Jul 30
ex-Google graph lead: in 4-6 months, every serious agent will run on a graph - no more one-file-at-a-time

Sunick shares an ex-Google graph lead's prediction: in 4-6 months, every serious agent will run on a graph, not one file at a time. Graph-based context (code, docs, and data linked as a graph) gives agents structured navigation instead of flat file reads. Bold timeline, interesting direction.

  • Prediction: agents move to graph-based context
  • Graphs give structured navigation over flat files
  • Watch whether the timeline holds
graphsagentsprediction
View original on X ↗

Best Kimi K3 explainer

Alex Lieberman (@businessbarista) · Jul 30
Best explainer on Kimi K3 i've read. It walks you through how the model works & the elegant innovations…

Alex Lieberman calls this the best Kimi K3 explainer he has read, walking through how the model works and its elegant innovations. Kimi K3 is a recurring star in your bookmarks; a genuinely good explainer is the fastest way to get current. Read this before the implementation posts.

  • Top-rated explainer of the Kimi K3 model
  • Read before diving into implementations
  • Elegant innovations explained clearly
kimiexplainermodels
View original on X ↗

How ChatGPT optimizes its agent loop

Bytebytego (@bytebytego) · Jul 29
How ChatGPT Optimizes its Agent Loop: Harness, API, and Inference

Bytebytego's article on how ChatGPT optimizes its agent loop across harness, API, and inference. The agent loop (perceive, decide, act, observe) has optimization opportunities at every layer. Seeing how the most-used agent product tunes its loop is practical education.

  • Agent loop optimized across harness, API, and inference
  • Every loop layer has tuning opportunities
  • Learn from the most-deployed agent product
chatgptagent-loopoptimization
View original on X ↗

NVIDIA: build agents as Python objects

elvis (@omarsar0) · Jul 29
Super interesting new work from NVIDIA. (bookmark it) They suggest building agents as Python objects…

Elvis shares new NVIDIA work suggesting agents be built as Python objects. Treating agents as objects (with state, methods, and composition) is a programming-model proposal: it makes agents feel like normal software. If NVIDIA is pushing it, expect tooling to follow.

  • Agents-as-Python-objects programming model
  • Makes agents feel like normal software components
  • NVIDIA backing suggests coming tooling
nvidiaagentsprogramming-model
View original on X ↗

Next-Latent Prediction

Jayden Teoh (@jayden_teoh_) · Jun 16
Next-token prediction is myopic. What if transformers learn to predict their own next latent state? We present Next-Latent Prediction (NextLat)…

Jayden Teoh presents Next-Latent Prediction (NextLat): what if transformers predict their own next latent state instead of the next token? Next-token prediction is called myopic; predicting latent states could capture longer-range structure. An intriguing architectural bet.

  • NextLat: predict latent states, not tokens
  • Challenges the myopia of next-token prediction
  • Architectural bet worth watching
architecturetransformersresearch
View original on X ↗

An AI agent is a while-loop

Alex Xu (@alexxubyte) · May 14
An AI agent can be thought of as a simple While-loop. It uses an LLM to select an action, executes…

Alex Xu's classic framing: an AI agent is a simple while-loop where an LLM selects an action, executes it, observes the result, and repeats. Stripping agents down to this loop demystifies them. Everything fancy (planning, memory, tools) is elaboration on this loop.

  • Agents reduced to: LLM picks action, executes, observes, repeats
  • Demystifies agent architecture
  • All advanced features elaborate on this loop
agentsmental-modelclassic
View original on X ↗

Block's plan to replace corporate hierarchy with AI

Rohan Paul (@rohanpaul_ai) · Mar 31
Jack Dorsey's Block just laid out a plan to replace much of corporate hierarchy with AI coordination…

Rohan Paul covers Jack Dorsey's Block laying out a plan to replace much of corporate hierarchy with AI coordination. It is a bold corporate vision: agents coordinating work that managers used to do. Read it as a signal of where tech leadership thinks this goes, not as a prediction to bank on.

  • Block's vision: AI coordination replacing management layers
  • Signal of tech leadership thinking, not a forecast
  • Provocative framing for org-design discussions
future-of-workblockai-coordination
View original on X ↗