📚 Learning

Courses, guides, handbooks, and deep dives. (73 entries)

Evals guide for AI agents in production

Alex Lieberman (@businessbarista) · Oct 5
long evals guide covering how to evaluate AI agents in production…

Alex Lieberman shares a long evals guide on evaluating AI agents in production. Production evals differ from benchmark evals: you measure task success, cost, latency, and failure modes on real traffic. This is core to your interests, so read it as a practitioner, not a tourist.

  • Guide to evaluating agents on real production traffic
  • Production evals cover success, cost, latency, failures
  • Read deeply; this is your core topic
evalsproductionagents
View original on X ↗

15 projects as an AI Evals Engineer

Suraj Sharma (@suraj_sharma14) · Oct 4
As an AI Evals Engineer… 15 projects…

Suraj Sharma frames 15 projects for someone becoming an AI Evals Engineer. Evals (systematic testing of AI behavior) is now a real specialty, and project lists like this sketch the skill surface: golden datasets, regression suites, human-vs-model agreement, and more. Useful as a curriculum checklist.

  • 15 projects mapping the AI evals engineer role
  • Use as a checklist of evals competencies
  • Pairs well with the 12-stage learning plan
evalsprojectscareer
View original on X ↗

6 months to become an AI Evals Engineer, 12 stages

Suraj Sharma (@suraj_sharma14) · Oct 3
If I had 6 months to become an AI Evals Engineer… 12 stages…

Suraj Sharma lays out 12 stages for becoming an AI Evals Engineer in 6 months. The stages likely run from statistics basics through golden datasets, LLM-as-judge, and production monitoring. As a staged plan it is more actionable than a flat list; the timeline is aspirational, the ordering is the value.

  • 12-stage plan toward AI evals engineering
  • The stage ordering matters more than the 6-month timeline
  • Cross-check stages against real job descriptions
evalslearning-plancareer
View original on X ↗

18-bullet AI learning list

Suraj Sharma (@suraj_sharma14) · Oct 2
CANCEL your weekend plans… 18 bullets…

Suraj Sharma's 'cancel your weekend plans' thread packs 18 learning bullets. These threads are best treated as menus, not assignments. Scan for the two or three items that fill your actual gaps (likely evals and inference, given your bookmarks) and ignore the rest.

  • 18-item learning menu in one thread
  • Treat as a menu: pick 2-3, ignore the rest
  • Match picks to your real gaps, not FOMO
learningcurated-list
View original on X ↗

LM text classification mega write-up

Sebastian Raschka (@rasbt) · Sep 29
mega write-up on LM text classification…

Sebastian Raschka's mega write-up on LM text classification. Raschka is known for rigorous, code-first explanations, and text classification with language models remains a bread-and-butter production task. Even in the agent era, classification is everywhere.

  • Raschka-quality write-up on LM classification
  • Classification remains a core production task
  • Code-first rigor you can trust
classificationnlpraschka
View original on X ↗

Agent Evals Handbook

Manish Sharma (@lucifer_x007) · Sep 25
Agent Evals Handbook: claude.ai/artifact/QeGEd…

Manish Sharma shares an Agent Evals Handbook as a Claude artifact. A handbook format suggests practical guidance: how to design evals, what to measure, how to avoid common traps. Evals are your stated interest area, so this deserves a real read, not a skim.

  • Practical handbook on agent evaluation
  • Likely covers eval design and common pitfalls
  • Read fully; this is your core topic
evalshandbookagents
View original on X ↗

AI Performance Engineering, curated

wafer (@wafer_ai) · Sep 19
AI Performance Engineering / Curated by Wafer

Wafer curates AI Performance Engineering resources on GitHub. Performance engineering for AI means making models run fast and cheap: kernels, batching, memory layout, hardware choice. A curated repo saves you from assembling the reading list yourself.

  • Curated performance-engineering resource collection
  • Covers the make-models-fast-and-cheap discipline
  • Start here before Googling scattered blog posts
performancegpucurated-list
View original on X ↗

Anthropic's 13-page PDF on Agent Memory

chuplung (@choopyplug1) · Sep 14
Anthropic just dropped a 13-page PDF on Agent Memory…

Chuplung shares Anthropic's 13-page PDF on Agent Memory. Memory (how agents remember across sessions: what to store, how to retrieve, when to forget) is one of the least-solved parts of agent design. Thirteen pages from Anthropic is the closest thing to an official reference.

  • Anthropic's official-style doc on agent memory
  • Memory is among the least-solved agent problems
  • Read as the reference, then compare with alternatives
memoryagentsanthropic
View original on X ↗

ai-abc: learning production AI

Suraj Sharma (@suraj_sharma14) · Sep 13
This is what learning production AI should've looked like

Suraj Sharma shares ai-abc.vercel.app with the claim 'this is what learning production AI should have looked like.' Production AI (models running in real products with real users) is a different discipline from notebook ML. If the site delivers structured production-first learning, it fills a genuine gap.

  • Production-first AI learning resource
  • Production AI differs sharply from notebook ML
  • Evaluate the site's structure before committing time
production-ailearningresource
View original on X ↗

MoBA breakdown

How To Prompt (@HowToPrompt__) · Sep 8
long MoBA breakdown…

How To Prompt shares a long MoBA breakdown. MoBA (Mixture of Block Attention) is an attention variant that routes computation to relevant blocks instead of attending to everything. New attention variants matter because attention dominates inference cost.

  • MoBA: block-routed attention variant explained
  • Attention variants target inference cost
  • Understand the routing idea, not just the acronym
attentionmobaarchitecture
View original on X ↗

Anthropic VLIW+SIMD homework

Zartbot (@zartbotF) · Sep 7
Tackling Anthropic's VLIW+SIMD homework…

Zartbot tackles Anthropic's VLIW+SIMD homework. VLIW (very long instruction word) and SIMD (single instruction, multiple data) are low-level parallelism techniques for squeezing performance from hardware. This is deep-systems content; valuable if you go deep on inference, skippable otherwise.

  • Deep dive into low-level parallelism (VLIW, SIMD)
  • Relevant only if you go deep on inference performance
  • Skim unless hardware-level work is your direction
systemssimdperformance
View original on X ↗

KV Caching in LLMs, Clearly Explained

Akshay (@akshay_pachaar) · Feb 9
No text captured (article or link save).

Akshay's article 'KV Caching in LLMs, Clearly Explained.' The KV cache stores computed key-value pairs during generation so the model does not recompute them for every new token; it is the single most important inference optimization. A clear explainer on this topic pays for itself many times over.

  • KV cache is the most important inference optimization
  • Clear explainer worth reading carefully
  • Foundation for understanding vLLM, SGLang, and serving
kv-cacheinferenceexplainer
View original on X ↗

12 Agentic AI Engineer projects

Suraj Sharma (@suraj_sharma14) · Sep 4
If you can build these 12 Agentic AI Engineer projects…

Suraj Sharma lists 12 Agentic AI Engineer projects to build. Project-based lists like this work because each project forces you through the full loop: design, build, evaluate, debug. Pick the ones closest to production agent work and skip the toy demos.

  • 12 projects aimed at agentic AI engineering
  • Learn by building end-to-end, not by reading
  • Prioritize the projects with real evaluation loops
projectsagentsportfolio
View original on X ↗

Build B200 attention kernel from scratch

Iaroslav Elistratov (@iaro_e) · Aug 27
Build B200 attention kernel from scratch in CUDA and a little PTX…

Iaroslav Elistratov walks through building a B200 attention kernel from scratch in CUDA with a little PTX. Kernels are the low-level GPU programs that make attention fast; the B200 is Nvidia's latest chip. Writing one yourself is the deepest possible way to learn how attention actually executes.

  • Write a real attention kernel for Nvidia's B200
  • PTX (assembly-level GPU code) included
  • The deepest way to learn attention execution
cudakernelsattention
View original on X ↗

Free CUDA course

Tech with Mak (@techNmak) · Aug 27
Nvidia should be paying this guy. This free course teaches CUDA…

Tech with Mak shares a free course teaching CUDA, joking that Nvidia should be paying the instructor. CUDA is Nvidia's GPU programming platform, and a free structured course removes the last excuse for not learning it. Pair with the 100 Days of CUDA journey for practice.

  • Free structured CUDA course
  • Removes the cost barrier to GPU programming
  • Pair coursework with a daily practice habit
cudacoursefree
View original on X ↗

Inference engineering archive, 0 to 100

Suraj Gaud (@notsurajgaud) · Aug 27
inference engineering archive, 0-100.

Suraj Gaud shares an inference engineering archive covering 0 to 100. Archives that span beginner to advanced are rare and valuable: you can enter at your level and keep going. Use it as the spine of an inference learning track.

  • Beginner-to-advanced inference archive
  • Use as the spine of your inference learning
  • Enter at your level, work upward
inferencearchivelearning
View original on X ↗

100 Days of CUDA journey

Pavel Simo (@pavelsimo) · Aug 24
i'm documenting my 100 Days of CUDA journey…

Pavel Simo documents a 100 Days of CUDA journey on his site. CUDA is Nvidia's programming platform for GPUs; learning it means learning how to think in parallel. A public 100-day log is useful both as a curriculum and as proof that sustained daily practice works.

  • 100-day public learning log for CUDA
  • Follow the progression as a ready-made curriculum
  • Daily public practice is itself a forcing function
cudagpulearning-journey
View original on X ↗

Learn LLM Inference Engineering step by step

Amit Shekhar (@amitiitbhu) · Aug 17
Learn LLM Inference Engineering step by step…

Amit Shekhar's GitHub guide to learning LLM inference engineering step by step. Inference engineering (serving models fast and cheap) is a distinct skill from training or prompting. A step-by-step repo is the right format for this hands-on discipline.

  • Step-by-step inference engineering guide
  • Serving models is its own discipline
  • Hands-on format suits the topic
inferencegithubtutorial
View original on X ↗

Inference Engineering, free book

Baseten (@baseten) · Aug 18
Inference Engineering is free for everyone to read: baseten.co/inference-engi

Baseten made their Inference Engineering book free for everyone to read. A full book on inference (serving, batching, quantization, hardware) from a company that does it professionally is a serious resource. Books beat threads for building complete mental models.

  • Complete inference engineering book, now free
  • From practitioners who serve models professionally
  • Read cover to cover for a complete mental model
inferencebookfree
View original on X ↗

Georgia Tech CS 8803 LLM class

Wei Xu (@cocoweixu) · Aug 17
Updated CS 8803 Large Language Model class at @GeorgiaTech…

Wei Xu shares the updated CS 8803 Large Language Model class at Georgia Tech. University courses bring structure and rigor that tweet threads lack: lectures, assignments, and a sensible ordering. Audit it like a real class for the topics you have not covered.

  • Full university LLM course, updated
  • Structure and rigor beat scattered threads
  • Audit the topics you have not covered properly
courseuniversityllm
View original on X ↗

LLM Engineer's Handbook

Akshay (@akshay_pachaar) · Aug 14
long LLM engineer's handbook…

Akshay's long LLM Engineer's Handbook on GitHub. Handbooks aim for completeness: the full surface of what an LLM engineer should know. Use it as a reference and gap-finder rather than a cover-to-cover read.

  • Comprehensive LLM engineer handbook
  • Use as reference and skill-gap finder
  • Not a cover-to-cover read; dip into gaps
handbookreferencellm
View original on X ↗

aiengineeringfromscratch.com

Suraj Sharma (@suraj_sharma14) · Aug 20
This is what learning AI should've looked like

Suraj Sharma shares aiengineeringfromscratch.com as 'what learning AI should have looked like.' From-scratch learning (implementing things yourself rather than calling APIs) builds the intuition that survives framework churn. Worth checking whether it delivers real implementations or just branding.

  • From-scratch AI learning site
  • Implementation-first learning builds durable intuition
  • Verify depth before investing serious time
learningfrom-scratchresource
View original on X ↗

10 inference concepts to implement yourself

TensorTonic (@TensorTonic) · Aug 14
10 inference concepts worth understanding by implementing them yourself…

TensorTonic lists 10 inference concepts worth understanding by implementing them yourself. Inference (running a trained model to get outputs) has a deep bag of tricks: batching, caching, quantization, and more. Implementing each one turns vague familiarity into real understanding.

  • 10 inference concepts, each learned by building
  • Covers the practical tricks behind fast model serving
  • Work through them in order for a complete picture
inferencehands-onlearning
View original on X ↗

AI Engineering Skills Map

Andrew Ng (@AndrewYNg) · Aug 14
New: A map of the most important skills in AI Engineering.

Andrew Ng publishes a map of the most important skills in AI Engineering. A skills map from Ng is useful as a compass: it shows which capabilities the field's educators consider durable. Compare it against your own skill inventory to find gaps worth closing.

  • Ng's map of durable AI engineering skills
  • Use it to audit your own skill coverage
  • Skills maps age well; revisit yearly
skillscareerandrew-ng
View original on X ↗

13 summer sessions on hamel.dev

Hamel Husain (@HamelHusain) · Aug 12
Over the summer, @sh_reya and I hosted 13 sessions…

Hamel Husain hosted 13 summer sessions, now on hamel.dev. Hamel's sessions are known for practical, no-hype applied LLM engineering. Thirteen sessions is a serious corpus; skim titles and watch the ones on evals and production patterns.

  • 13 practical sessions on applied LLM engineering
  • Known for honest, production-grounded content
  • Prioritize evals and production topics
sessionsapplied-llmhamel
View original on X ↗

Lecture 75: ScaleML

Sayak Paul (@RisingSayak) · Aug 6
No text captured (article or link save).

Sayak Paul shares Lecture 75 on ScaleML via YouTube. Lecture-numbered content suggests a real course series on scaling machine learning. Late-numbered lectures usually assume background; check earlier lectures if the prerequisites feel shaky.

  • Lecture 75 in a ScaleML series
  • Late lectures assume background; backfill as needed
  • Part of a larger structured course
scalinglecturevideo
View original on X ↗

Inside vLLM

pdawg (@prathamgrv) · Aug 5
No text captured (article or link save).

Pdawg shares 'Inside vLLM' from aleksagordic.com. vLLM is the dominant open-source LLM serving engine; understanding its internals (paged attention, continuous batching) explains why it is fast. Reading internals turns a black-box tool into something you can tune and debug.

  • vLLM internals: paged attention, continuous batching
  • Understanding internals lets you tune and debug serving
  • Essential if vLLM is in your stack
vllminferenceinternals
View original on X ↗

20-video free post-training course

Nathan Lambert (@natolambert) · Aug 7
My summer project is done! A 20 video, free course on post-training to accompany my book is all on YouTube…

Nathan Lambert finished a 20-video free course on post-training to accompany his book, all on YouTube. Post-training (everything after pretraining: instruction tuning, RLHF, DPO) is where models become useful products. Twenty videos from a practitioner who wrote the book is a serious free resource.

  • 20 free videos on post-training, companion to his book
  • Covers instruction tuning through RLHF and beyond
  • One of the best free structured resources in your pile
post-trainingcoursefree
View original on X ↗

15 ML projects

Suraj Sharma (@suraj_sharma14) · Aug 6
CANCEL your weekend plans. You NEED to: 15 ML projects…

Suraj Sharma's weekend list: 15 ML projects to build. Like his other lists, the value is in doing, not collecting. Fifteen is too many; pick three that stretch different muscles (data, modeling, deployment).

  • 15 ML project ideas in one thread
  • Pick 3 that cover different skills, ignore the rest
  • Doing beats bookmarking
projectsmlweekend
View original on X ↗

11-step SAE walkthrough

Tom Yeh (@ProfTomYeh) · Aug 6
11-step SAE walkthrough…

Tom Yeh shares an 11-step walkthrough of SAEs (sparse autoencoders), a technique for interpreting what neural networks actually learned by finding the sparse features inside them. SAEs are a hot interpretability tool. A step-by-step walkthrough is the gentlest on-ramp.

  • 11 steps through sparse autoencoders
  • SAEs reveal the features a model actually uses
  • Good entry to the interpretability literature
interpretabilitysaetutorial
View original on X ↗

inferenceengineering.tech

Shubham Mishra (@shubh6200) · Aug 5
Learning inference from @philipkiely at inferenceengineering.tech If you want to actually understand LLM serving, KV caching, and model infra instead of just prompting wrappers, bookmark this site immediately. Easily one of the best, most practical engineering resources out right now.

Shubham Mishra recommends inferenceengineering.tech by Philip Kiely as one of the best, most practical engineering resources for understanding LLM serving, KV caching, and model infra. Practical and engineering-focused beats theoretical for this topic. Bookmark the site itself, not just the tweet.

  • Highly recommended practical inference resource
  • Covers serving, KV caching, and model infra
  • Bookmark the site, work through it systematically
inferenceresourcepractical
View original on X ↗

18 AI Engineer projects

Suraj Sharma (@suraj_sharma14) · Aug 4
CANCEL your weekend plans. You NEED to: 18 AI Engineer projects…

Another Suraj Sharma weekend list: 18 AI Engineer projects. The pattern across his lists is consistent, build in public, ship complete things. If you have done none, start with one; if you have done many, these are portfolio filler, not learning.

  • 18 AI engineer project ideas
  • One finished project beats 18 bookmarked lists
  • Best for early-career portfolio building
projectsai-engineeringportfolio
View original on X ↗

21-item learn list

Mohit Goyal (@ByteMohit) · Aug 3
21-item learn list…

Mohit Goyal shares a 21-item learn list. Long lists are starting points, not curricula. Without knowing your level, the honest advice: use it to spot topics you have never touched, then find one good resource per gap.

  • 21 topics in one list
  • Use it to find blind spots, not as a syllabus
  • One good resource per gap beats 21 shallow ones
learningcurated-list
View original on X ↗

50 LLM interview questions

Tech with Mak (@techNmak) · Aug 4
50 LLM interview questions…

Tech with Mak posts 50 LLM interview questions. Given your active job search, this is directly useful: interview loops for AI roles now probe LLM internals, evals, and agent design, not just ML basics. Work through them as a mock-interview bank.

  • 50 LLM interview questions in one place
  • Directly useful for your current job search
  • Practice answering out loud, not just reading
interviewsllmcareer
View original on X ↗

Microsoft 24-lesson AI curriculum

Tech with Mak (@techNmak) · Aug 2
Learn AI From the Ground Up. This is a gem by Microsoft, a 24-lesson curriculum that actually covers the fundamentals. It's currently #1 trending on GitHub today…

Tech with Mak shares Microsoft's 24-lesson AI curriculum, calling it a gem that covers fundamentals and noting it trended #1 on GitHub. Microsoft's open curricula are consistently solid. Twenty-four lessons is a real course; treat it like one.

  • 24-lesson fundamentals curriculum from Microsoft
  • Trended #1 on GitHub, signaling quality
  • Treat as a real course with exercises
curriculummicrosoftfundamentals
View original on X ↗

SVD: the math behind Netflix, Google, PCA

Mathematica (@mathemetica) · Aug 2
The singular value decomposition that quietly powers Netflix, Google, and PCA…

Mathematica's thread on singular value decomposition: the math quietly powering Netflix recommendations, Google search, and PCA. SVD factors a matrix into meaningful components; it is one of those ideas that keeps reappearing. Worth understanding once, properly.

  • SVD powers recommendations, search, and PCA
  • One of the most reused ideas in applied math
  • Learn it once, recognize it forever
mathsvdfundamentals
View original on X ↗

Read the ACT paper, implement from scratch

Keivalya Pandya (@KeivalyaP) · Aug 2
CANCEL your weekend plans. You NEED to: Read the ACT paper. Then implement it from scratch…

Keivalya Pandya's challenge: read the ACT paper, then implement it from scratch. ACT (Adaptive Computation Time) lets models spend more compute on harder inputs. Paper-to-implementation is the highest-value learning exercise in ML.

  • Read ACT paper, implement from scratch
  • Paper-to-code is the highest-value ML exercise
  • Adaptive compute is an important efficiency idea
paperimplementationact
View original on X ↗

Every resource shared on X, in one place

Avinash Singh (@AvinashSingh_20) · Jul 16
Every Resource I Have Ever Shared on X, All in One Place. Over the past few months I have shared a lot of resources on X. Company databases, visa sponsorship…

Avinash Singh compiles every resource he has ever shared on X into one place, including company databases and visa sponsorship info. Mega-compilations are useful as a searchable attic. Skim once, then search it when you need something specific.

  • Mega-compilation of shared resources
  • Use as a searchable attic, not a reading list
  • Includes practical items like visa sponsorship data
resourcescompilationreference
View original on X ↗

LLM internals step by step

Amit Shekhar (@amitiitbhu) · Aug 2
Learn LLM internals step by step - from tokenization to attention to inference optimization…

Amit Shekhar walks through LLM internals step by step: tokenization, attention, inference optimization. Internals-first learning builds intuition that survives API changes. This is the right order: tokens in, math in the middle, optimized serving at the end.

  • Tokenization to attention to inference optimization
  • Internals-first learning survives framework churn
  • Follow the steps in order
internalsllmtutorial
View original on X ↗

Day 26/90 of Inference Engineering

max fu (@maxxfuu) · Aug 2
Day 26/90 of Inference Engineering I didn't understand global memory coalescing until I took a step back…

Max Fu's day 26 of 90 on inference engineering covers global memory coalescing: arranging memory access so the GPU reads efficiently. Memory access patterns are often the real bottleneck, not raw compute. The series format (90 days) is a solid learning commitment device.

  • Memory coalescing: the hidden inference bottleneck
  • 90-day series format keeps momentum
  • Memory patterns matter more than raw FLOPs
inferencegpumemory
View original on X ↗

Best transformers explanation video

Wayne (@waynoir) · Aug 1
This hour long video is genuinely the best explanation of transformers on the internet…

Wayne shares an hour-long video called the best transformers explanation on the internet. Transformers are the architecture behind all modern LLMs. An hour is a real investment; the payoff is a mental model that makes every paper easier to read.

  • Hour-long deep explainer on transformers
  • Builds the mental model that unlocks papers
  • Worth the hour if transformers still feel fuzzy
transformersvideoexplainer
View original on X ↗

Day 210/365 of GPU Programming

levi (@levidiamode) · Aug 1
Day 210/365 of GPU Programming Reviewing differences in tensor core data layouts between Nvidia hardware…

Levi's day 210 of a 365-day GPU programming series covers tensor core data layouts across Nvidia hardware generations. Tensor cores are the specialized units that make AI math fast; their data layouts (how numbers are arranged in memory) change between chip generations. Deep, specific, and genuinely useful for kernel work.

  • Tensor core layouts across Nvidia generations
  • Deep specifics for real kernel optimization
  • Follow the series if GPU programming is your track
gputensor-corescuda
View original on X ↗

Sizing a production LLM deployment

h100envy (@h100envy) · Jul 31
NVIDIA Senior Data Scientist explained how to size a production LLM deployment in 33 minutes…

An NVIDIA senior data scientist explains how to size a production LLM deployment in 33 minutes of video. Sizing means choosing the right GPUs, counts, and configuration for your traffic. Practical capacity planning is rare content; most resources stop at the model.

  • 33-minute guide to sizing LLM deployments
  • Covers GPUs, counts, and configuration for real traffic
  • Rare practical capacity-planning content
deploymentsizingvideonvidia
View original on X ↗

Laplace transforms and Mamba

fintex (@_yusufknl) · Jul 31
As someone who ships LLM systems in production, this 22-minute Laplace transform video is the closest thing to watching Mamba get invented in real time…

Fintex shares a 22-minute Laplace transform video, calling it the closest thing to watching Mamba get invented in real time. Laplace transforms are the math behind continuous-time systems; Mamba (a state-space model) builds on related ideas. The video connects classical math to a modern architecture.

  • Laplace transforms connected to Mamba's design
  • Classical math behind a modern architecture
  • Watch for the 'how was this invented' insight
mathmambavideo
View original on X ↗

TamilLM from scratch: the book

Ramchand Kumaresan (@Mechramc) · Jul 30
Everything I needed to learn to build TamilLM from scratch is in this book. Tokenizers, attention, KV-cache, RoPE, GQA, MoE, RLHF…

Ramchand Kumaresan shares the book behind building TamilLM from scratch: tokenizers, attention, KV cache, RoPE, GQA, MoE, RLHF. A from-scratch book covering the full modern stack in one place is rare. Even if you never build a Tamil model, the stack walkthrough is the education.

  • Full modern LLM stack in one from-scratch book
  • Tokenizers through RLHF, all implemented
  • The stack walkthrough is the real value
from-scratchbookllm-stack
View original on X ↗

40-minute PyTorch tutorial

AVB (@neural_avb) · Jul 29
Here is a 40 minute Pytorch tutorial that explains its main mechanics step-by-step. Lot of breadth covered

AVB shares a 40-minute PyTorch tutorial covering the main mechanics step by step with breadth. PyTorch is the framework almost all AI research and much production runs on. Forty focused minutes on mechanics (tensors, autograd, modules) is a solid refresher or first pass.

  • 40 minutes on PyTorch mechanics, step by step
  • Good refresher on tensors, autograd, and modules
  • Breadth-first: follow up on whatever felt shaky
pytorchtutorialvideo
View original on X ↗

12 YouTube videos to become an AI engineer

Hasan Toor (@hasantoxr) · Jul 30
If you want to become a world-class AI engineer in 2026, watch these 12 YouTube videos (save them):

Hasan Toor lists 12 YouTube videos to watch to become a world-class AI engineer in 2026. Curated video lists save the search time but vary wildly in quality. Treat it as a starting playlist, then go deeper on the topics that matter for your work.

  • 12-video curated playlist for AI engineering
  • Use as a starting point, not a complete education
  • Go deeper on evals, harness, and inference topics
videoscurated-listlearning
View original on X ↗

Day 7/30: Chunked Prefill

prasanna (@jaga_prasanna) · Jul 29
Day 7/30 of inference infrastructure series Chunked Prefill what happens when a massive prompt hits your inference server…

Prasanna's day 7 of a 30-day inference infrastructure series explains chunked prefill: what happens when a massive prompt hits your inference server. Prefill (processing the input prompt) vs decode (generating tokens) is the fundamental split in LLM serving. Chunking large prefills keeps the server responsive.

  • Chunked prefill keeps servers responsive on huge prompts
  • Prefill vs decode is the core serving distinction
  • Follow the 30-day series for the full picture
inferencevllmserving
View original on X ↗

Lilian Weng's RL blogs

Mohammed Alshehri (@M0EGPT) · Jul 28
Lilian was my favourite guide into RL. Her blogs were the first things I read when I wanted to dive deeper…

Mohammed Alshehri recommends Lilian Weng's RL blogs as his favorite guide into reinforcement learning. Lilian Weng's blog is legendary for clear, thorough explanations of RL and LLM topics. Start with her RL posts, then follow into her LLM content.

  • Lilian Weng's blog: the classic RL on-ramp
  • Clear, thorough, and mathematically honest
  • Start with RL posts, continue to LLM content
reinforcement-learningblogclassic
View original on X ↗

vLLM + SGLang for high-throughput inference

Suraj Sharma (@suraj_sharma14) · Jul 28
CANCEL your weekend plans. You NEED to: Learn vLLM + SGLang for high-throughput inference. Build…

Suraj Sharma says to learn vLLM and SGLang for high-throughput inference. These are the two leading open-source serving engines; knowing both (and their tradeoffs) is table stakes for inference engineering. Learn one deeply, survey the other.

  • vLLM and SGLang are the two serving engines to know
  • Learn one deeply, survey the other
  • Table stakes for inference engineering
vllmsglanginference
View original on X ↗

Kimi K3 paper, implemented step by step

Deep-ML (@real_deep_ml) · Jul 28
took a while but got it done, you can now understand the Kimi K3 paper by implementing it step by step

Deep-ML implemented the Kimi K3 paper step by step so readers can understand it by following along. Paper-to-implementation walkthroughs are the best way to truly understand an architecture. Given how often Kimi K3 appears in your bookmarks, this is a high-priority read.

  • Kimi K3 paper implemented step by step
  • Understand the architecture by following the code
  • High priority given your Kimi K3 interest
kimiimplementationpaper
View original on X ↗

How to scale your model, full guide

Matt Dancho (@mdancho84) · Jul 27
Replying to @mdancho84 FULL GUIDE: HOW TO SCALE YOUR MODEL…

Matt Dancho's full guide on how to scale your model, with JAX links. Scaling (making models bigger and training on more data efficiently) involves parallelism strategies, memory planning, and infrastructure. JAX-focused, so most relevant if you work in that ecosystem.

  • Full guide to model scaling with JAX
  • Covers parallelism and memory planning
  • Ecosystem-specific; adapt concepts to your stack
scalingjaxtraining
View original on X ↗

48 hours with the Kimi K3 modeling code

ali (@waterloo_intern) · Jul 27
I spent 48 hours with the Kimi K3 modeling code. It took: - 650 mg of caffeine (mandatory) - 40 cans…

Ali spent 48 hours with the Kimi K3 modeling code, fueled by caffeine, and reports back on video. Reading a real model implementation for two days straight is an education in how architectures actually fit together. The video format lets you follow the reasoning, not just the conclusions.

  • 48-hour deep read of Kimi K3's modeling code
  • Code reading reveals how architectures really fit together
  • Follow along with the repo open
kimicode-readingvideo
View original on X ↗

Golden dataset design for AI regression testing

Suraj Sharma (@suraj_sharma14) · Jul 26
CANCEL your weekend plans. You NEED to: Learn golden dataset design for regression testing AI systems…

Suraj Sharma covers golden dataset design for regression testing AI systems. A golden dataset is a curated set of input-output pairs you re-run after every change to catch regressions. This is the unglamorous backbone of production AI quality, and most teams do it badly.

  • Golden datasets are the backbone of AI regression testing
  • Design them carefully: coverage, difficulty, freshness
  • Most teams underinvest here; it is a differentiator
evalsdatasetstesting
View original on X ↗

12 GenAI projects: build to get hired

Suraj Sharma (@suraj_sharma14) · Jul 26
If you can build these 12 GenAI projects. You're hired. Project 1: Production API Wrapper FastAPI…

Suraj Sharma claims that building 12 GenAI projects gets you hired, starting with a production FastAPI wrapper. The framing is motivational rather than literal, but the underlying point is sound: shipped projects beat certificates. A FastAPI wrapper around a model is genuinely the right first project.

  • 12 projects framed as a hiring signal
  • Project 1 (production API wrapper) is the right starting point
  • Shipped work beats credentials in this market
projectscareergenai
View original on X ↗

Structured LLM learning path

Zeng Jiajun (@zengjiajun_eth) · Jul 25
if @karpathy's from zero to hero series doesn't fit your taste, and you wanna learn llm in a more structured/academic way…

Zeng Jiajun offers a structured, academic alternative to Karpathy's Zero to Hero for learning LLMs. Karpathy's series is beloved but informal; a structured path suits learners who want rigor and ordering. Pick the style that matches how you learn.

  • Academic-structured alternative to Karpathy's series
  • Choose based on your learning style
  • Structure helps if self-direction is hard
learning-pathllmstructured
View original on X ↗

Modular LLM Inference Handbook

Junu Park (@junupark_) · Jul 26
Replying to @junupark_ handbook.modular.com LLM Inference

Junu Park shares Modular's LLM Inference Handbook. Modular (the Mojo company) works on next-generation AI infrastructure, so their handbook reflects a systems-builder perspective. Compare its advice against the vLLM-centric mainstream.

  • Inference handbook from Modular's systems perspective
  • Compare against vLLM-centric mainstream advice
  • Systems-builder viewpoint adds depth
inferencehandbookmodular
View original on X ↗

YouTube playlist

Tech with Mak (@techNmak) · Jul 25
Replying to @techNmak Here's the Playlist: youtube.com

Tech with Mak shares a YouTube playlist as a reply. The playlist link is the content; context suggests it complements his other learning threads. Playlists are only as good as their curation, so skim the video titles before committing.

  • Companion YouTube playlist
  • Skim titles before committing watch time
  • Playlists need curation to be worth it
playlistvideolearning
View original on X ↗

Optimizers in Deep Learning, by building them

TensorTonic (@TensorTonic) · Jul 25
Optimizers in Deep Learning, Explained by Building Them

TensorTonic's article: optimizers in deep learning, explained by building them. Optimizers (the algorithms that update model weights during training, like Adam) are usually treated as black boxes. Building them yourself reveals what the hyperparameters actually do.

  • Build optimizers to understand them
  • Demystifies Adam and friends
  • Hyperparameters make sense after implementation
optimizersdeep-learninghands-on
View original on X ↗

Fourier series lecture and attention

fintex (@_yusufknl) · Jul 24
As someone who reads every transformer paper that drops, this Fourier series lecture is the closest thing to watching attention get invented…

Fintex shares a Fourier series lecture described as the closest thing to watching attention get invented. Fourier series decompose signals into frequencies; the connection to attention is about how models represent and mix information. Beautiful math, and it deepens intuition for why attention works.

  • Fourier series lecture with a surprising link to attention
  • Builds mathematical intuition for representation mixing
  • Watch for depth, not immediate practicality
mathfourierattention
View original on X ↗

Andrew Ng's 8-page PDF: 4 agentic steps

codila (@0xCodila) · Jul 23
Andrew Ng just dropped 8-page PDF on 4 agentic steps from Loops to Graphs from scratch…

Codila shares Andrew Ng's 8-page PDF on 4 agentic steps, from loops to graphs, built from scratch. Ng's gift is compression: the agentic design space reduced to four patterns you can actually implement. Eight pages is a lunch-break read with lasting value.

  • 4 agentic patterns from loops to graphs, in 8 pages
  • Implement each pattern to lock in the understanding
  • Ng's compression of the design space is the value
agentsandrew-ngpatterns
View original on X ↗

Agent concepts, made simple

ai maxxer (@agentsmaxxing) · Jul 23
This guy makes complex agent concepts surprisingly easy to understand. If you're serious about agents…

An AI maxxer account recommends a creator who makes complex agent concepts surprisingly easy to understand, via video. Good explainers are worth more than dense papers when you are building intuition. Check whether the simplification preserves the important details.

  • Recommended explainer for agent concepts
  • Good for building intuition before papers
  • Verify simplifications against primary sources
agentsexplainervideo
View original on X ↗

30-problem PyTorch sheet

pdawg (@prathamgrv) · Jul 22
while learning MLSys and inference, i built a 30 pytorch problem sheet at @TensorTonic it includes popular topics like - MLA, GQA, MHA, MoE - quantization, memory related stuff - kv cache

Pdawg built a 30-problem PyTorch problem sheet covering MLA, GQA, MHA, MoE, quantization, memory, and KV cache. Problem sheets beat passive reading: each problem forces you to produce the answer. This one targets exactly the topics that matter for inference engineering interviews and real work.

  • 30 problems on attention variants, MoE, quantization, KV cache
  • Active problem-solving beats passive reading
  • Doubles as interview prep for inference roles
pytorchproblemsinference
View original on X ↗

LLM Cache Management resources

Praveen Kumar Verma (@Alacritic_Super) · Jul 21
Want to master LLM Cache Management? Start with these resources. KV Cache…

Praveen Kumar Verma collects resources on mastering LLM cache management, starting with KV cache. Cache management (what to keep, what to evict, how to share across requests) is where inference efficiency is won or lost. A resource list on this narrow topic is genuinely useful.

  • Focused resource list on LLM cache management
  • Caching is where inference efficiency is decided
  • Narrow topic, high practical value
kv-cacheinferenceresources
View original on X ↗

Unsloth: RL, kernels, reasoning in 2 hours

h100envy (@h100envy) · Jul 9
Ex-NVIDIA engineer who built Unsloth explained RL, kernels, reasoning, quantization, and agents in 2 hours…

h100envy shares a 2-hour video where the ex-NVIDIA engineer behind Unsloth explains RL, kernels, reasoning, quantization, and agents. Unsloth made fine-tuning dramatically faster; learning from its builder compresses years of experience. Two hours is a real session; take notes.

  • Unsloth builder on RL, kernels, quantization, agents
  • 2-hour deep session from a practitioner
  • Take notes; density is high
unslothfine-tuningvideo
View original on X ↗

Study the NVIDIA Hopper architecture

Praveen Kumar Verma (@Alacritic_Super) · Jul 9
If you want to master LLM training and inference, study the NVIDIA Hopper Architecture…

Praveen Kumar Verma advises studying the NVIDIA Hopper architecture to master LLM training and inference. Hopper (the H100 generation) introduced transformer-specific hardware features; understanding the hardware explains why software is shaped the way it is. Hardware literacy is underrated in AI engineering.

  • Hopper architecture explains modern AI software design
  • Hardware literacy is underrated
  • Know the chip to understand the stack
hardwarenvidiahopper
View original on X ↗

Attention is a lookup

tetsuo (@tetsuoai) · Jul 2
Attention is a lookup. Each token builds a query, compares it against every key in the sequence, and…

Tetsuo explains attention as a lookup: each token builds a query, compares it against every key in the sequence, and pulls a weighted mix of values. Attention is the core operation of transformers, and the lookup framing is the clearest mental model. Short video, high value per minute.

  • Attention explained as query-key-value lookup
  • The clearest mental model for the transformer's core op
  • Watch before any deeper transformer content
attentiontransformersexplainer
View original on X ↗

277-page LLM secrets PDF

Matt Dancho (@mdancho84) · Jun 11
This 277-page PDF unlocks the secrets of Large Language Models. Here's what's inside:

Matt Dancho shares a 277-page PDF on the secrets of large language models. Long PDFs are reference material, not reads: skim the table of contents, read the chapters that fill your gaps, keep the rest for lookup. Check the date, since LLM knowledge goes stale fast.

  • 277-page reference PDF on LLMs
  • Read as a reference: TOC first, then gap chapters
  • Check recency; LLM content ages quickly
pdfreferencellm
View original on X ↗

Claude Code power-user thread

Rohit (@rohit4verse) · Mar 30
I thought I was a Claude Code power user. Then the creator of Claude Code (@bcherny) dropped this thread and I realized I've been using maybe 30% of the tool…

Rohit thought he was a Claude Code power user until the creator's thread showed him he was using maybe 30% of the tool. The thread curates the creator's tips. The meta-lesson: with fast-moving tools, periodically re-learn from the source instead of coasting on old habits.

  • Curated thread of Claude Code power tips
  • Even experienced users were at 30% utilization
  • Re-learn fast-moving tools from the source periodically
claude-codeproductivitytools
View original on X ↗

Claude Code hidden features, by its creator

Boris Cherny (@bcherny) · Mar 29
I wanted to share a bunch of my favorite hidden and under-utilized features in Claude Code. I'll focus…

Boris Cherny, the creator of Claude Code, shares his favorite hidden and under-utilized features. Learning a tool from its creator beats learning from tutorials: you get the intended workflows, not folk workarounds. Focus on the features that change how you structure work, not just shortcuts.

  • Hidden features straight from Claude Code's creator
  • Creator workflows beat tutorial folk knowledge
  • Adopt the structural features, not just shortcuts
claude-codeproductivitytools
View original on X ↗

Claude How-To visual guide

Charlie Hills (@charliejhills) · Mar 27
BREAKING: Someone built a visual guide to master Claude Code in a weekend. It's called Claude How-To…

Charlie Hills shares Claude How-To, a visual guide to mastering Claude Code in a weekend. Visual guides work well for tools with spatial workflows (panes, modes, context). A weekend is enough to go from dabbling to fluent if you actually build something during it.

  • Visual guide to Claude Code mastery
  • Build a real thing during the weekend for it to stick
  • Good onboarding ramp for the tool
claude-codeguidevisual
View original on X ↗

Build an AI agent today, full course

hoeem (@hooeem) · Mar 26
I want to build an AI agent today (full course): No-one has made a full course so that anyone (yes, you) can create an AI agent from scratch…

Hoeem shares a full course article: build an AI agent today, aimed at anyone. Full courses that promise 'anyone can do it' vary in depth, but a complete build-along is still useful for seeing the whole loop once. Best for getting the end-to-end shape before going deep.

  • Full course on building an AI agent from scratch
  • Good for seeing the end-to-end loop once
  • Follow with deeper material on evals and harness
courseagentsbeginner
View original on X ↗

Andrew Ng's free AI agents course of 2026

Kanika (@KanikaBK) · Mar 19
JUST IN: 7 hours ago ANDREW NG dropped one of the MOST VALUABLE FREE AI COURSES OF 2026. Most AI agents… Save now!

Kanika flags Andrew Ng's free AI agents course, called one of the most valuable free AI courses of 2026. Ng's courses are consistently the best-structured free option for a topic. If you take one course from this whole pile, make it this one.

  • Ng's free agentic AI course, highly recommended
  • Best-structured free option in your bookmarks
  • Pair it with the 8-page PDF for theory plus practice
courseagentsandrew-ngfree
View original on X ↗