🔨 Build ideas

Repos, tools, and techniques worth trying or building on. (25 entries)

RL environments for training agents

clem (@ClementDelangue) · Oct 5
long post on RL environments for training agents…

Clem Delangue (Hugging Face CEO) shares RL environments for training AI agents, pointing to Hugging Face's FineEnv spaces. Reinforcement learning, where an agent learns by trial and error against a reward signal, needs good sandboxed environments to practice in. These spaces give builders ready-made training grounds instead of building their own from scratch.

  • Hugging Face is hosting ready-made RL environments (FineEnv) for agent training
  • Good environments are the bottleneck for RL-trained agents, not just algorithms
  • Worth exploring as practice grounds for any agent harness work
reinforcement-learningagent-traininghuggingface
View original on X ↗

388 AI agent skills

beamnxw (@beamnxw) · Sep 27
388 AI agent skills

A collection of 388 AI agent skills, shared with a video walkthrough. Agent skills are reusable capability packs (tools, prompts, and know-how bundled together) that an agent can pick up for specific jobs. A library this large is worth mining for patterns: which skills are actually useful, and how they are packaged.

  • 388 reusable agent skills in one collection, a pattern library for skill design
  • Study how skills are packaged: tools plus prompts plus domain know-how
  • Candidate source of ideas for your own agent skill packs
agent-skillstoolingpatterns
View original on X ↗

LangExtract by Google

Matt Dancho (@mdancho84) · Sep 24
RIP document extractors. Google just released LangExtract…

Google released LangExtract, a tool for pulling structured data out of documents, and Matt Dancho declares traditional document extractors dead. Instead of brittle regex or template parsing, you describe what you want and the model extracts it into clean structured output. For anyone processing PDFs, contracts, or reports, this is a tool to test head-to-head against current pipelines.

  • LangExtract turns documents into structured data using an LLM instead of templates
  • Test it against your existing extraction pipelines before rebuilding anything
  • Signals Google pushing hard into the document-AI space
document-aiextractiongoogle
View original on X ↗

6.12B-request LLM inference traces dataset

Juncheng Yang (@1a1a11a) · Sep 19
Announcing one year of LLM inference metadata traces, 6.12B requests…

Juncheng Yang announces one full year of LLM inference metadata traces, 6.12 billion requests, published as an open dataset. Inference metadata (timings, batch sizes, queue depths, hardware counters) is gold for anyone studying how models actually serve in production. This is rare public data for evals and systems research.

  • 6.12B inference requests of metadata, openly published
  • Useful for studying real serving behavior: batching, latency, queueing
  • Strong dataset for evals or inference-engineering research
datasetinferenceresearch
View original on X ↗

Personal agent x Okta integration pattern

Marc Klingen (@marcklingen) · Sep 8
Looking forward to seeing a personal agent <> Okta integration…

Marc Klingen looks forward to a personal agent integrating with Okta, the identity and access management platform companies use for single sign-on. The interesting idea: an agent that can act with your corporate identity, calling internal tools as you. That raises exactly the governed-access questions your own approval-gate work addresses.

  • Personal agents plus corporate identity (Okta) is the enterprise agent pattern
  • Acting-as-you agents need scoped permissions and approval gates
  • Directly relevant to your governed agent runtime thinking
identityenterprise-agentsokta
View original on X ↗

Reef: model and harness co-improvement infra

Bo Liu (@Benjamin_eecs) · Sep 1
Check out Reef: open-source infra for model and harness co-improvement at test time:))

Bo Liu shares Reef, open-source infrastructure for model and harness co-improvement at test time. The idea: instead of only improving the model or only the harness (the surrounding code that runs the agent), improve both together while the system runs. This sits very close to your RRSI interest in self-improving agent harnesses.

  • Reef co-improves model and harness together at test time
  • Open source, so the mechanism is inspectable, not just claimed
  • Strong overlap with RRSI-style self-improvement research
self-improvementagent-harnessopen-source
View original on X ↗

TangleML self-improving loops infra

tobi lutke (@tobi) · Sep 1
Btw we open sourced the core infra piece that makes these self improving loops possible.

Shopify CEO Tobi Lutke announces open-sourcing the core infrastructure behind self-improving loops at TangleML. Self-improving loops are systems where an agent's outputs feed back into making the next run better, automatically. A production deployment of this idea, from a serious engineering org, is worth studying closely.

  • TangleML's self-improving loop infra is now open source
  • Production-proven feedback loops, not a research demo
  • Compare its loop design against RRSI-style approaches
self-improvementopen-sourceproduction
View original on X ↗

AMD free API

Samgor (@biggor888) · Sep 1
AMD's free API is here: developer.amd.com.cn/radeon/tokenfa (translated from Chinese)

AMD opened a free API endpoint for developers (via developer.amd.com.cn). Free inference APIs are useful for prototyping agents without paying per-token fees during development. Worth bookmarking as a cost-free backend for experiments, with the usual caveat that free tiers can change or rate-limit.

  • Free AMD API endpoint for developers
  • Useful as a zero-cost backend for agent prototypes
  • Watch for rate limits and terms changes on free tiers
apifree-tierprototyping
View original on X ↗

35 production-grade agentic AI architectures

Tom Doerr (@tom_doerr) · Aug 30
Implements 35 production-grade agentic AI architectures…

Tom Doerr shares a GitHub repo implementing 35 production-grade agentic AI architectures. Having that many architectures side by side is a comparative goldmine: planners, routers, multi-agent setups, and more, all in one place. Skim the list first, then deep-dive the two or three closest to what you are building.

  • 35 agentic architectures implemented in one repo
  • Use it as a comparative catalog before committing to a design
  • Good reference when you need 'how do others structure this'
architecturesgithubreference
View original on X ↗

NAC by Arcee AI

Lucas Atkins (@latkins) · Aug 13
NAC: github.com

Lucas Atkins shares NAC from Arcee AI (github.com/arcee-ai/nac). The tweet itself is terse, just the name and repo, so the substance is in the repository. Worth a look at what Arcee (a small-model specialist) considers worth open-sourcing.

  • NAC is an Arcee AI open-source release
  • Context is thin; the repo README is the real content
  • Arcee focuses on small efficient models, so expect that flavor
open-sourcearceerepo
View original on X ↗

Search stack: Google + Perplexity + Exa

Mrinal (@Hi_Mrinal) · Aug 15
Bookmarked it 4 days ago…With Google, perplexity and exa…

Mrinal shares a bookmarked search setup combining Google, Perplexity, and Exa. Each covers a different retrieval style: Google for breadth, Perplexity for answered synthesis, Exa for semantic and neural search over the web. For research-heavy agent work, a multi-engine retrieval stack beats any single source.

  • Combine Google, Perplexity, and Exa for complementary retrieval
  • Different engines cover breadth, synthesis, and semantic search
  • A pattern to copy for agent research pipelines
searchretrievalresearch
View original on X ↗

Graphify: 107K-star repo

Safi (@safishamsii) · Aug 18
107,000+ GitHub stars. 5.2M+ downloads… (@graphify)

Safi highlights Graphify: 107,000+ GitHub stars and 5.2M+ downloads. Whatever it does, that adoption curve is worth understanding. Star counts that high usually mean it solved a painful, common problem simply. Look at what it is and why it spread.

  • 107K stars signals a widely-felt pain point solved well
  • Study the README and issue tracker to learn what users love
  • Adoption mechanics are as interesting as the tech
open-sourcegithubtrending
View original on X ↗

dimos

Swastika Yadav (@swstica) · Aug 20
No text captured (article or link save).

Swastika Yadav shares dimos (github.com/dimensionalOS/dimos) with no accompanying text. The repository is the content here. Given the naming (dimensional OS), it likely relates to agent operating-system-style infrastructure. Worth a direct look at the repo.

  • Bare repo link; the README carries the substance
  • Name suggests OS-like infrastructure for agents
  • Verify what it actually does before investing time
open-sourcerepoinfrastructure
View original on X ↗

DeepSeek V4 Pro free access

James Grugett (@jahooma) · Aug 13
We're offering DeepSeek V4 Pro 100% free…

James Grugett announces DeepSeek V4 Pro available 100% free via freebuff.com. Free access to a frontier-class model is useful for evals, distillation experiments, or just heavy prototyping without a bill. As always with free model endpoints, check data handling before sending anything sensitive.

  • DeepSeek V4 Pro free through freebuff.com
  • Good for evals and prototyping without inference costs
  • Do not send sensitive data to free third-party endpoints
free-tierdeepseekmodels
View original on X ↗

Kimi Delta Attention, minimal implementation

Arjun (@arjunkocher) · Aug 7
Kimi Delta Attention. - a minimal implementation of the Kimi K3 architecture.

Arjun shares a minimal implementation of Kimi Delta Attention, the attention mechanism from the Kimi K3 architecture. Attention is the core operation where each token weighs every other token; Kimi's variant is a novel twist. A minimal implementation is the fastest way to understand what actually changed.

  • Minimal implementation of Kimi K3's attention mechanism
  • Reading small code beats reading the paper for intuition
  • Kimi K3 is a recurring theme in your bookmarks; this is the hands-on entry
kimiattentionimplementation
View original on X ↗

Microsoft Skill Recorder, open source

AI SuperDomain (@AISuperDomain) · Aug 1
Microsoft has open-sourced a very interesting project: Skill Recorder…

Microsoft open-sourced Skill Recorder, a project for recording agent skills. If skills are reusable capability packs, a recorder is the tooling that captures a human or agent demonstration and turns it into a reusable skill. This is infrastructure for the skill-based agent economy.

  • Microsoft's open-source tool for recording agent skills
  • Demonstration-to-skill is a key authoring workflow
  • Watch how big labs think skills should be created
agent-skillsmicrosoftopen-source
View original on X ↗

GEPA for fine-tuning on a budget

Gajesh (@gajesh) · Jul 26
if you're broke trying to do fine-tuning - pivot to GEPA; Imagine…

Gajesh suggests pivoting to GEPA when fine-tuning budgets run dry. GEPA (likely an efficient adaptation method) promises fine-tuning-like gains without fine-tuning-scale compute. For anyone who has priced out a full fine-tune and flinched, efficient alternatives are worth knowing.

  • GEPA as a budget alternative to full fine-tuning
  • Efficient adaptation methods matter when compute money is tight
  • Verify the claims on your own workload before committing
fine-tuningefficiencygepa
View original on X ↗

10 must-try open source repos

Farhan Azad Shuvra (@Faazsh) · May 20
10 must-try open source repos for the weekend: 1. Clawless…

Farhan Azad Shuvra lists 10 must-try open source repos for the weekend, starting with Clawless. Weekend-sized repo lists are good for breadth: clone two or three, run them, keep what sticks. Clawless at the top suggests agent-adjacent tooling.

  • Curated weekend list of 10 repos
  • Start with Clawless and expand from there
  • Time-box exploration: run, evaluate, keep or drop
open-sourcecurated-listweekend-projects
View original on X ↗

Goose: free Claude Code alternative

Nav Toor (@heynavtoor) · Apr 5
Claude Code costs $200/month. GitHub Copilot costs $19/month. Jack Dorsey's company built a free alternative. 35,000 GitHub stars. It's called Goose…

Nav Toor compares prices: Claude Code at $200/month, Copilot at $19/month, then points to Goose, a free alternative with 35,000 GitHub stars from Jack Dorsey's company (Block). An agentic coding assistant that is free and open source changes the build-vs-buy math for coding agents.

  • Goose: free, open-source agentic coding assistant, 35K stars
  • From Block (Jack Dorsey's company), so serious backing
  • Evaluate it as the free baseline before paying for coding agents
coding-agentsopen-sourcegoose
View original on X ↗

Forked Claude Code running on any LLM

Twigpine (@twigpine) · Mar 31
We forked the leaked Claude Code source and made it work with ANY LLM: GPT, DeepSeek, Gemini, Llama,…

Twigpine shares a fork of the leaked Claude Code source, modified to work with any LLM: GPT, DeepSeek, Gemini, Llama. Decoupling a strong agent harness from its original model is exactly the kind of experiment that reveals how much of the magic is harness vs model. Legally and practically murky given the leak origin, but technically fascinating.

  • Claude Code harness ported to run on any model
  • Tests the harness-vs-model contribution question directly
  • Note the leaked-source origin before building on it
claude-codeagent-harnessopen-source
View original on X ↗

claw-code: Claude Code rewritten in Python

Gergely Orosz (@GergelyOrosz) · Mar 31
Replying to @GergelyOrosz The repo: github.com/… The brilliance: copyright does not protect derived works. Rewriting TypeScript code in Python means…

Gergely Orosz covers claw-code: Claude Code rewritten from TypeScript to Python, arguing copyright does not protect the derived reimplementation. A Python rewrite makes the harness hackable for the Python-native AI crowd. The legal reasoning is as interesting as the code.

  • Claude Code reimplemented in Python as claw-code
  • Python version is far more hackable for most AI engineers
  • The copyright-derivation argument is worth understanding
claude-codepythonopen-source
View original on X ↗

Kelly Grants: $100K in AI credits

kelly criterion (@thekellycrit) · Aug 10
Announcing Kelly Grants: I'm gifting up to $100,000 in AI usage credits & compute…

Kelly (kelly criterion) announces Kelly Grants: gifting up to $100,000 in AI usage credits and compute. Free credits are runway for experiments, evals, or training runs you would not otherwise fund. Grant programs like this are worth applying to early, before the pool is picked over.

  • Up to $100K in AI credits and compute being gifted
  • Apply early; grant pools get competitive fast
  • Free compute extends your experiment runway
grantscomputefunding
View original on X ↗

Qwen3.8-27B in BF16 on free Kaggle TPU

Md Ismail Sojal (@0x0SojalSec) · Sep 6
You can now run Qwen3.8-27B in full BF16 on a free Kaggle TPU…

A post shows running Qwen3.8-27B in full BF16 precision on a free Kaggle TPU. BF16 (a 16-bit number format) halves memory vs full precision while keeping training-stable range. Free TPU hours plus a capable open model is a strong combo for fine-tuning and inference experiments at zero hardware cost.

  • 27B model in BF16 fits on free Kaggle TPU
  • Free TPUs are viable for serious open-model experiments
  • Good setup for fine-tuning trials without GPU bills
qwentpukagglefree-tier
View original on X ↗

GMI Cloud: 14 days free GPU

GMI Cloud (@gmi_cloud) · Aug 24
unlimited MiniMax M3 and M2.7, 14 days FREE on GMI Cloud…

GMI Cloud offers 14 days free on MiniMax M3 and M2.7 GPUs. Two free weeks of modern GPU time is enough for a real eval run, a fine-tune, or inference benchmarking. Worth using for a bounded experiment with a clear question, not idle tinkering.

  • 14 days of free MiniMax GPU time
  • Enough for a real eval or fine-tuning experiment
  • Go in with a specific question to make the window count
gpufree-tiercloud
View original on X ↗

Z.ai GLM-5.3 free

painn (@painn_x) · Aug 19
long post on Z.ai GLM-5.3 FREE…

Z.ai's GLM-5.3 is available free, per this post. GLM models (from Zhipu) are strong open-weight contenders, and free access makes them easy to evaluate against your current stack. Add it to the model comparison set for agent backends.

  • GLM-5.3 available free via Z.ai
  • Another strong open model for agent backend comparisons
  • Benchmark it on your evals before adopting
glmfree-tiermodels
View original on X ↗