Essays, breakdowns, and explainers worth one focused read. (47 entries)
Linear for agents: tools, evals, harness design
Aishwarya Goel (@aishwarya_08) · Oct 5
Aishwarya Goel's long thread, 'Linear for agents,' covering tools for AI agents: evals, harness design, and production patterns. Linear (the project tool) is invoked as an analogy for the kind of tight, well-designed tooling agents need. A good survey of the agent tooling landscape as of October 2026.
- Survey of AI agent tools: evals, harness, production
- 'Linear for agents' sets a design-quality bar
- Good landscape overview, recent
agentstoolingevals
View original on X ↗
Stop arguing about which AI agent is best
helicerat (@helicerat0x) · Oct 5
Helicerat argues people should stop debating which AI agent is best, with images and a Claude Cowork mention. The point: the harness, evals, and workflow matter more than the model badge. Tool debates are usually proxy wars for workflow differences.
- Agent debates miss the point; harness and workflow matter more
- Claude Cowork as the current reference
- Focus on your loop, not the leaderboard
agentsopiniontooling
View original on X ↗
Uber MCP infrastructure
santi (@santtiagom_) · Oct 3
Santi shares a translated post on Uber's MCP infrastructure. MCP (Model Context Protocol) is the emerging standard for connecting agents to tools and data. Seeing how Uber builds MCP infrastructure at scale is a preview of enterprise agent plumbing.
- Uber-scale MCP infrastructure breakdown
- MCP is becoming the agent-tool standard
- Enterprise plumbing preview
mcpuberinfrastructure
View original on X ↗
Bad code and AGENTS.md
Tomas Vykruta (@tvykruta) · Sep 29
Tomas Vykruta's long thread on bad code and AGENTS.md. AGENTS.md is the emerging convention of a markdown file that tells AI coding agents how to work in your repo. The thread connects code quality to agent effectiveness: agents amplify whatever codebase you give them, good or bad.
- AGENTS.md: instructions for AI agents in your repo
- Agents amplify your codebase's quality, good or bad
- Clean code matters more, not less, in the agent era
agents-mdcode-qualitycoding-agents
View original on X ↗
System 1 vs System 2 Agent Harnesses
Avi Chawla (@_avichawla) · Sep 28
Avi Chawla frames agent harnesses as System 1 vs System 2, borrowing the fast/slow thinking distinction. System 1 harnesses react quickly with cached patterns; System 2 harnesses reason deliberately. The framing helps decide when an agent should think fast vs think hard.
- Fast vs slow thinking applied to harness design
- Match harness depth to task difficulty
- Useful framing for your own harness decisions
harnessagentsmental-model
View original on X ↗
DoorDash Vera breakdown
Yarchi (@undefinedKi) · Sep 27
Yarchi breaks down DoorDash's Vera, the internal data agent serving 10,000+ employees: curated context, enforced access, real evals, and plan approval. Vera went from 43% to 90% answer accuracy by beating plain Claude Code by 25 points with the same model. You have studied this deeply already; it remains the canonical enterprise agent case study.
- Vera: 43% to 90% accuracy via context engineering, not a better model
- The four pillars: curated context, enforced access, real evals, plan approval
- Canonical case study for enterprise agent architecture
veradoordashenterprise-agentscase-study
View original on X ↗
How to Design an Agent Harness
Yarchi (@undefinedKi) · Aug 15
Yarchi's article: six decisions that turn a model into a worker you can leave alone. The harness decisions (tools, memory, evals, recovery, permissions, observability) are the difference between a demo and a system. Six decisions is a manageable framework for designing your own.
- Six harness decisions for autonomous workers
- Framework for designing your own harness
- Directly applicable to your agent runtime work
harnessagentsdesign
View original on X ↗
Xiaomi MiMo V2.6 RL open weights
Adithya S K (@adithya_s_k) · Sep 26
Adithya S K shares Xiaomi's MiMo V2.6 RL model on Hugging Face. Open-weight RL-trained models are valuable for research and for builders who want to inspect or fine-tune. Another data point in the open-model wave.
- MiMo V2.6 RL open weights on Hugging Face
- Useful for research and fine-tuning
- Part of the accelerating open-model wave
open-weightsxiaomimodels
View original on X ↗
Ontology is at the center
Chad Wahlquist (@chadwahl) · Sep 24
Chad Wahlquist argues ontology is at the center, linking Palantir. An ontology (a structured map of what things exist in a domain and how they relate) is what lets AI systems reason about a business instead of just predicting text. This connects directly to your Bloomberg-ontology and Palantir MMDP studies.
- Ontology as the center of enterprise AI
- Structured domain maps enable reasoning, not just prediction
- Connects to your Bloomberg and Palantir research
ontologypalantirenterprise-ai
View original on X ↗
Halo now supports SDPG
Yifan Zhang (@yifanzhang_) · Sep 21
Yifan Zhang announces Halo now supports SDPG, linking a paper (arxiv 2606.04036). SDPG appears to be a training or alignment method; Halo is the framework adopting it. New method support in a popular framework is how research becomes usable.
- Halo adds support for the SDPG method
- Framework adoption is how research becomes usable
- Check the paper for what SDPG actually does
trainingmethodshalo
View original on X ↗
Loop Engineering for Sub-Dime Agents
Google Cloud Tech (@GoogleCloudTech) · Sep 21
Google Cloud Tech's article on Loop Engineering for Sub-Dime Agents: making agent loops so cheap each run costs under ten cents. Cost per task is the gating factor for agent deployment; sub-dime loops change what products are viable. Engineering for cost is as important as engineering for capability.
- Engineering agent loops down to sub-dime cost
- Cost per task gates real deployment
- Cheap loops enable new product categories
agentscostengineering
View original on X ↗
SpaceX AI engineer: 2,500 PRs a month
kloess (@kloss_xyz) · Sep 21
Kloess shares the story of a SpaceX AI engineer shipping 2,500 PRs a month. The number is startling and probably reflects heavy agent assistance. It is a glimpse of what AI-accelerated engineering velocity looks like at the extreme.
- 2,500 PRs/month: extreme AI-assisted velocity
- Glimpse of the agent-accelerated future
- Ask what breaks at that speed, not just what ships
productivityagentsspacex
View original on X ↗
Safety theater
CYBERDOG (@oCYBERDOGo) · Sep 18
CYBERDOG's long post on safety theater: safety measures that look reassuring but do not actually reduce risk. The concept matters because theater consumes the budget and attention that real safety work needs. Read critically; the term gets overused, but the underlying point is real.
- Safety theater: reassuring measures that do not reduce risk
- Theater crowds out real safety work
- Apply the lens skeptically, including to this post
ai-safetycritiqueessay
View original on X ↗
Agent Substrate: K8s infrastructure
Ayaan (@twtayaan) · Sep 17
Ayaan writes a long post on Agent Substrate: Kubernetes infrastructure for agents. Kubernetes (K8s) orchestrates containers at scale; an agent substrate is the platform layer agents run on. As agents move to production, the infrastructure questions (scheduling, isolation, scaling) become as important as the model questions.
- Agent substrate: the platform layer agents run on
- K8s brings scheduling, isolation, and scaling to agents
- Infrastructure matters as much as models in production
infrastructurekubernetesagents
View original on X ↗
What OpenAI's agentic software factory looks like
Gergely Orosz (@GergelyOrosz) · Sep 15
Gergely Orosz reveals what OpenAI's agentic software factory looks like. An agentic software factory is OpenAI's internal system for producing software with AI agents at scale. A rare look inside how the frontier lab industrializes agent-assisted coding.
- Inside OpenAI's agent-powered software production
- How the frontier lab industrializes coding agents
- Compare with Uber's software factory piece
openaiagentsengineering
View original on X ↗
RSI algorithms date back to 1987
Juergen Schmidhuber (@SchmidhuberAI) · Sep 13
Juergen Schmidhuber notes concrete algorithms for recursive self-improvement date back to 1987. RSI (systems that improve themselves) is often discussed as a future breakthrough; Schmidhuber points out the formal foundations are decades old. Essential historical context for your RRSI work.
- RSI formalisms date to 1987, not the future
- Historical context for self-improvement research
- Schmidhuber's priority claims are worth knowing
rsihistoryself-improvement
View original on X ↗
Closed-loop and cybernetics
Yi Ma (@YiMaTweets) · Sep 13
Yi Ma's long post on closed-loop systems and cybernetics, linking to lab materials. Cybernetics is the old science of control and communication in machines and animals; closed-loop means the system's outputs feed back into its inputs. These ideas predate modern AI and quietly underpin RL and agent design.
- Cybernetics: the science of feedback and control
- Closed-loop thinking underpins RL and agents
- Old ideas, newly relevant
cyberneticscontrol-theoryfoundations
View original on X ↗
ToolGrad from Google Research
Google Research (@GoogleResearch) · Sep 10
Google Research introduces ToolGrad. The name suggests gradient-like optimization applied to tool use: improving how agents select and use tools through a learning signal. New agent-training methods from Google Research deserve attention.
- ToolGrad: optimizing agent tool use
- From Google Research, so take it seriously
- Read the linked material for the mechanism
googleagentsresearch
View original on X ↗
AI safety orgs list
Owain Evans (@OwainEvans_UK) · Sep 9
Owain Evans shares a list of AI safety organizations. The safety ecosystem is fragmented across research labs, nonprofits, and company teams; a map of who is who saves a lot of confusion. Useful context whether or not safety is your focus.
- Map of the AI safety organization landscape
- Saves confusion in a fragmented ecosystem
- Good context even for non-safety work
ai-safetyorganizationsreference
View original on X ↗
Tiny models
Ryan Wright (@IQReactorAI) · Sep 3
Ryan Wright's long post on tiny models. Small models are getting capable enough for real tasks at a fraction of the cost and latency. The tiny-model trend matters for agents: cheap, fast models change what you can afford to put in a loop.
- Tiny models are becoming genuinely useful
- Cost and latency advantages compound in agent loops
- Watch the capability-per-dollar curve
small-modelsefficiencyagents
View original on X ↗
Stuck in a massive local optimum
Tim Rocktaschel (@_rockt) · Sep 3
Tim Rocktaschel argues the field is stuck in a massive local optimum. A local optimum is a good-but-not-best solution that is hard to leave because every small step looks worse. The provocation: current approaches work well enough to discourage the exploration that could find something better.
- Are we stuck in a local optimum?
- Good-enough methods discourage exploration
- Provocation worth arguing with
opinionresearchprovocation
View original on X ↗
How to RL post-train a 397B model
Edward Hu (@edwardjhu) · Sep 1
Edward Hu asks how one RL post-trains a 397B model for long-horizon knowledge work. Post-training a 397-billion-parameter model with reinforcement learning is at the edge of what is feasible; the infrastructure and algorithmic challenges are extreme. Read for the scale of the problem, not a how-to.
- RL post-training at 397B parameters
- Extreme infrastructure and algorithmic challenges
- Read for perspective on the scaling frontier
post-trainingscalerl
View original on X ↗
Inference-stack rant
Nick Khami (@skeptrune) · Aug 31
Nick Khami's long inference-stack rant. Rants from practitioners often contain more truth than polished posts: the frustrations reveal where the tooling genuinely hurts. Read for the pain points; they are product and research opportunities.
- Practitioner frustrations with the inference stack
- Pain points are product opportunities
- Rants reveal truth that polished posts hide
inferenceopinionrant
View original on X ↗
Three links: LangChain, Kaggle, Claude blogs
Manuel Zapata (@ManuelZapata) · Aug 29
Manuel Zapata shares three links: a LangChain blog on agents, a Kaggle whitepaper, and a Claude blog post. Three substantive reads in one tweet. The Claude blog post is likely the highest signal; LangChain's blog tracks the framework's evolution.
- Three agent reads: LangChain, Kaggle, Claude
- Prioritize the Claude blog post
- Kaggle whitepaper for the data angle
agentsreading-listlinks
View original on X ↗
Apple Agent Seer paper
elvis (@omarsar0) · Aug 29
Elvis covers Apple's Agent Seer paper in a long post. Apple publishes sparingly on AI, so their agent research gets extra attention. Worth reading for how a privacy-first, on-device company thinks about agents differently from cloud labs.
- Apple's take on agent research
- Privacy-first, on-device perspective differs from cloud labs
- Rare Apple AI publication; read closely
appleagentspaper
View original on X ↗
Agentic Kernels in Production
Brian Li (@BrianLi23) · Aug 28
Brian Li's article on Agentic Kernels in Production. Kernels here likely means the core execution units of agent systems running reliably in production. Production concerns (reliability, observability, cost control) separate real systems from demos.
- Agentic kernels running in production settings
- Production concerns: reliability, observability, cost
- Separates real systems from demos
agentsproductionkernels
View original on X ↗
Running a Software Factory Efficiently at Uber Scale
Uber Engineering (@UberEng) · Aug 28
Uber Engineering's article on running a software factory efficiently at Uber scale. A 'software factory' is the industrialized pipeline for producing software: CI/CD, testing, deployment at massive scale. Uber-scale lessons on developer productivity and system reliability.
- Uber-scale software production practices
- Industrialized development pipelines
- Lessons on reliability at scale
engineeringuberscale
View original on X ↗
GLM 5.3 scaling with post-training, explained
Alex Ker (@thealexker) · Aug 28
Alex Ker's article explains GLM 5.3's scaling with post-training, intuitively. Post-training scaling (getting more from training after pretraining) is where much current progress comes from. An intuitive explainer beats the paper for first contact.
- Intuitive explainer of GLM 5.3's post-training
- Post-training scaling is a key progress lever
- Read before the paper
glmpost-trainingexplainer
View original on X ↗
Google wiki skills paper
DAIR.AI (@dair_ai) · Aug 28
DAIR.AI covers Google's wiki skills paper. Wiki skills likely refers to agents learning skills from wiki-style knowledge. Google's agent research output is prolific; this one is worth a skim for the skill-acquisition angle.
- Google paper on wiki-derived skills
- Skill acquisition is a key agent frontier
- Skim for the core idea
googleskillspaper
View original on X ↗
Meta/Stanford paper breakdown
marfin (@marfinxx) · Aug 26
Marfin's long post breaking down a Meta/Stanford paper. Joint industry-academic papers often combine scale with rigor. The breakdown format saves reading the full paper if you only need the key results.
- Meta/Stanford joint paper explained
- Industry scale plus academic rigor
- Read the breakdown, then decide on the paper
papermetastanford
View original on X ↗
Cursor on reliable Git hosting
Cursor (@cursor_ai) · Aug 18
Cursor writes about making Git hosting more reliable, performant, and scalable. Version control infrastructure is invisible until it breaks; Cursor's scale makes their lessons valuable. Relevant if you run infrastructure, skimmable otherwise.
- Reliable Git hosting at Cursor's scale
- Invisible infra until it breaks
- Skimmable unless you run infra
gitinfrastructurecursor
View original on X ↗
Thoughts about scaling laws
jietang (@jietang) · Aug 19
Jietang's long post sharing thoughts about scaling laws. Scaling laws (predictable relationships between compute, data, model size, and performance) underpin billion-dollar training decisions. Thoughtful commentary helps separate what is established from what is hype.
- Commentary on scaling law science
- Separates established results from hype
- Useful framing for training discussions
scaling-lawstheoryessay
View original on X ↗
DeepSeek harness talk summary
Phosphen (@phosphenq) · Aug 16
Phosphen summarizes a DeepSeek harness talk on video. DeepSeek's engineering culture produces unusually frank technical talks. Harness talks from strong labs are high-signal for your interests.
- DeepSeek harness talk summarized
- Strong labs give frank technical talks
- High signal for harness design
deepseekharnesstalk
View original on X ↗
Irreducible error essay
helicerat (@helicerat0x) · Aug 14
Helicerat's long essay on irreducible error. Irreducible error is the portion of mistakes no model can eliminate, the noise floor of a task. Understanding it calibrates expectations: some failures are not fixable with better models, only with better task design.
- Irreducible error: the mistakes no model can fix
- Calibrates expectations for agent reliability
- Some failures need task redesign, not better models
errortheoryessay
View original on X ↗
Update on our post-training effort
Gabe Pereyra (@gabepereyra) · Aug 20
Gabe Pereyra's article with an update on a post-training effort. Post-training (instruction tuning, RLHF, and friends) is where raw models become useful products. Practitioner updates reveal what actually worked, which is rarer than theory.
- Practitioner update on post-training work
- What actually worked, not just theory
- Compare methods against your own needs
post-trainingpractitionerupdate
View original on X ↗
Microsoft AI Frontiers internship project
Xidulu (@xidulu) · Aug 11
Xidulu shares a project from a Microsoft AI Frontiers internship, with a paper link. Internship projects at frontier labs are often surprisingly substantial. Worth a look at what the next generation of researchers is working on.
- Microsoft AI Frontiers internship project
- Frontier-lab intern work is often substantial
- Window into emerging research directions
researchmicrosoftinternship
View original on X ↗
Long essay on Canada
Eliot Pence (@EliotPence) · Aug 6
Eliot Pence's long essay on Canada. Topic is outside your usual AI bookmarks, which suggests it caught your eye for personal reasons. Long essays reward slow reading; skim the opening to decide if it earns the full read.
- Long-form essay on Canada
- Outside your usual topics; read for interest
- Skim the opening to judge the full read
essaycanada
View original on X ↗
5 papers to read this year
ajay yadav (@bettersayAJ) · Aug 2
Ajay Yadav says if you only read 5 more papers this year, make it these: the GPT-3 paper (few-shot learning), CLIP (connecting images and text), RAG (retrieval-augmented generation), and two more. These are foundation papers; reading them is rite-of-passage material for the field.
- 5 foundation papers: GPT-3, CLIP, RAG, and more
- Rite-of-passage reading for the field
- Read the papers, not just summaries
papersfoundationsreading-list
View original on X ↗
Verifiability
Yarchi (@undefinedKi) · Aug 1
Yarchi's long post on verifiability. Verifiability means being able to check an agent's work independently, not just trust its outputs. For enterprise agents, verifiable outputs are the difference between a demo and something you can deploy with accountability.
- Verifiable agent outputs enable real accountability
- Trust-but-verify beats blind trust in production
- Core requirement for enterprise deployment
verifiabilityagentstrust
View original on X ↗
Harness engineering as a discipline
beamnxw (@beamnxw) · Jul 30
Beamnxw shares a computer science paper establishing harness engineering as a proper engineering discipline for agents. The harness (everything around the model: tools, prompts, evals, loops) is where production agent quality actually lives. A paper arguing it deserves discipline-status validates your whole build direction.
- Paper argues harness engineering is its own discipline
- The harness, not the model, is where quality lives
- Academic validation of your build direction
harnessagentspaper
View original on X ↗
Every serious agent will run on a graph
sunick (@Serantych) · Jul 30
Sunick shares an ex-Google graph lead's prediction: in 4-6 months, every serious agent will run on a graph, not one file at a time. Graph-based context (code, docs, and data linked as a graph) gives agents structured navigation instead of flat file reads. Bold timeline, interesting direction.
- Prediction: agents move to graph-based context
- Graphs give structured navigation over flat files
- Watch whether the timeline holds
graphsagentsprediction
View original on X ↗
Best Kimi K3 explainer
Alex Lieberman (@businessbarista) · Jul 30
Alex Lieberman calls this the best Kimi K3 explainer he has read, walking through how the model works and its elegant innovations. Kimi K3 is a recurring star in your bookmarks; a genuinely good explainer is the fastest way to get current. Read this before the implementation posts.
- Top-rated explainer of the Kimi K3 model
- Read before diving into implementations
- Elegant innovations explained clearly
kimiexplainermodels
View original on X ↗
How ChatGPT optimizes its agent loop
Bytebytego (@bytebytego) · Jul 29
Bytebytego's article on how ChatGPT optimizes its agent loop across harness, API, and inference. The agent loop (perceive, decide, act, observe) has optimization opportunities at every layer. Seeing how the most-used agent product tunes its loop is practical education.
- Agent loop optimized across harness, API, and inference
- Every loop layer has tuning opportunities
- Learn from the most-deployed agent product
chatgptagent-loopoptimization
View original on X ↗
NVIDIA: build agents as Python objects
elvis (@omarsar0) · Jul 29
Elvis shares new NVIDIA work suggesting agents be built as Python objects. Treating agents as objects (with state, methods, and composition) is a programming-model proposal: it makes agents feel like normal software. If NVIDIA is pushing it, expect tooling to follow.
- Agents-as-Python-objects programming model
- Makes agents feel like normal software components
- NVIDIA backing suggests coming tooling
nvidiaagentsprogramming-model
View original on X ↗
Next-Latent Prediction
Jayden Teoh (@jayden_teoh_) · Jun 16
Jayden Teoh presents Next-Latent Prediction (NextLat): what if transformers predict their own next latent state instead of the next token? Next-token prediction is called myopic; predicting latent states could capture longer-range structure. An intriguing architectural bet.
- NextLat: predict latent states, not tokens
- Challenges the myopia of next-token prediction
- Architectural bet worth watching
architecturetransformersresearch
View original on X ↗
An AI agent is a while-loop
Alex Xu (@alexxubyte) · May 14
Alex Xu's classic framing: an AI agent is a simple while-loop where an LLM selects an action, executes it, observes the result, and repeats. Stripping agents down to this loop demystifies them. Everything fancy (planning, memory, tools) is elaboration on this loop.
- Agents reduced to: LLM picks action, executes, observes, repeats
- Demystifies agent architecture
- All advanced features elaborate on this loop
agentsmental-modelclassic
View original on X ↗
Block's plan to replace corporate hierarchy with AI
Rohan Paul (@rohanpaul_ai) · Mar 31
Rohan Paul covers Jack Dorsey's Block laying out a plan to replace much of corporate hierarchy with AI coordination. It is a bold corporate vision: agents coordinating work that managers used to do. Read it as a signal of where tech leadership thinks this goes, not as a prediction to bank on.
- Block's vision: AI coordination replacing management layers
- Signal of tech leadership thinking, not a forecast
- Provocative framing for org-design discussions
future-of-workblockai-coordination
View original on X ↗