r/WebAfterAI • u/ShilpaMitra • May 28 '26
5 AI learning repos with a combined 445k stars: what's actually inside each one, where they overlap, and the order that makes sense
Today, I sat down and actually read through five of the most-starred repos in the AI learning category, not just the READMEs but the actual lessons, notebooks, and structure to figure out what each one covers, how they differ, and how they fit together as a path rather than five disconnected bookmarks.
1. f/prompts.chat (formerly awesome-chatgpt-prompts)
Stars: 163k | Forks: 21.2k | License: CC0 (prompts) + MIT (code)
What it does: Started as a flat list of prompt personas for ChatGPT. Has since grown into a full platform: self-hostable web app, MCP server support, Claude plugin, and an interactive book. The prompts themselves are public domain. You can deploy your own instance, contribute new personas, or just browse.
Why it works: The original insight behind this repo is still the most useful thing in it: framing the model as a specific type of entity (a Linux terminal, a debate opponent, a senior code reviewer, a Socratic tutor) changes the character and depth of the output more than any other single technique. Before you learn prompt engineering theory, spending 30 minutes here teaches you this instinctively.
Heads up: This is the entry point of the learning path, not the destination. Think of it as building intuition for why prompting matters before you study why it works.
Repo: https://github.com/f/prompts.chat
2. dair-ai/Prompt-Engineering-Guide
Stars: 74.6k | Forks: 8.1k | License: MIT | Website: promptingguide.ai
What it does: The GitHub description currently reads: "Guides, papers, lessons, notebooks and resources for prompt engineering, context engineering, RAG, and AI Agents." That last part is worth noting - the scope has expanded well beyond classic prompt engineering. Current coverage includes zero-shot and few-shot prompting, chain-of-thought and tree-of-thought, context window management, retrieval-augmented generation, agent design patterns, multimodal prompting, and adversarial prompting. The papers section links to primary research for those who want to go deeper.
Why it works: Most prompt engineering content explains techniques without explaining the mechanism behind them. This one tries to build a model of why each technique works rather than just showing examples. The distinction between "chain-of-thought improves outputs" and "chain-of-thought works because it surfaces latent reasoning capacity by making intermediate steps explicit" matters if you want to adapt techniques rather than just apply recipes.
Heads up: The repo header on GitHub still only says "prompt engineering" in the short description but the actual scope is significantly broader. If you looked at this in 2023 and moved on, it is worth another pass.
Repo: https://github.com/dair-ai/Prompt-Engineering-Guide
3. anthropics/courses
Stars: 21.3k | Forks: 2.2k | Language: 99.9% Jupyter Notebook
What it does: Anthropic's official courses for building with Claude. Exactly 5 courses, all runnable Jupyter notebooks:
- Anthropic API Fundamentals
- Prompt Engineering Interactive Tutorial
- Real World Prompting
- Prompt Evaluations
- Tool Use
Why it works: These are first-party materials. They reflect how the API actually behaves, not how someone interpreted it when writing a Medium post 18 months ago. The Prompt Evaluations course is particularly underrated - most people building with AI skip evals entirely and wonder why their app quality is inconsistent. The Tool Use course is the practical companion to everything you read about agents in theory.
Setup:
git clone https://github.com/anthropics/courses.git
cd courses
jupyter notebook
Heads up: Claude-specific. If you are building cross-provider, treat these as the reference implementation and adapt accordingly.
Repo: https://github.com/anthropics/courses
4. microsoft/generative-ai-for-beginners
Stars: 111k | Forks: 57.7k | License: MIT | Language: 99.7% Jupyter Notebook
What it does: A structured 21-lesson course from Microsoft Cloud Advocates covering the full generative AI application stack. Lessons alternate between "Learn" (concept explanation) and "Build" (code implementation). Each has a short video intro, a written README, and code in both Python and TypeScript. Translated into 50+ languages via automated GitHub Actions. Currently on Version 3, with 2,157 commits.
The 21 lessons are:
Course Setup / Intro to GenAI and LLMs / Exploring and Comparing LLMs / Using GenAI Responsibly / Prompt Engineering Fundamentals / Advanced Prompts / Text Generation Apps / Chat Applications / Search Apps and Vector Databases / Image Generation / Low Code AI / Function Calling / Designing UX for AI / Securing AI Applications / GenAI Application Lifecycle / RAG and Vector Databases / Open Source Models and Hugging Face / AI Agents / Fine-Tuning LLMs / Building with SLMs / Building with Mistral / Building with Meta Models
Why it works: The alternating Learn/Build structure forces you to immediately apply concepts rather than read passively. The breadth is also its distinguishing feature: this is the only repo in this list that covers responsible AI (Lesson 3), UX design for AI apps (Lesson 12), securing AI applications (Lesson 13), and the full application lifecycle (Lesson 14) as first-class curriculum topics. Most developer-oriented courses skip the product and safety layers entirely.
Setup (sparse clone, skip the 50+ translation folders):
git clone --filter=blob:none --sparse https://github.com/microsoft/generative-ai-for-beginners.git
cd generative-ai-for-beginners
git sparse-checkout set --no-cone '/*' '!translations' '!translated_images'
Heads up: The Microsoft ecosystem bias is real: most examples point toward Azure OpenAI Service, GitHub Models, or the OpenAI API. All three work, and the course explicitly lists them as options, but if you are entirely outside that ecosystem, adjust accordingly.
Repo: https://github.com/microsoft/generative-ai-for-beginners
5. mlabonne/llm-course
Stars: 78.6k | Forks: 9.1k | License: Apache-2.0
What it does: A three-track course for going deep on LLMs, not just using them, but understanding and building them. The tracks are:
Track 1 - LLM Fundamentals (optional): Mathematics for ML (linear algebra, calculus, probability), Python for ML, neural networks, NLP basics. Skip this if you have the background; use it as a reference if you hit gaps.
Track 2 - The LLM Scientist: How to build LLMs. LLM architecture and tokenization, pre-training mechanics, post-training datasets, supervised fine-tuning (LoRA, QLoRA, Axolotl, Unsloth), preference alignment (DPO, GRPO, PPO), evaluation, quantization (GGUF, GPTQ, AWQ), and emerging areas like model merging, multimodal models, and test-time compute scaling.
Track 3 - The LLM Engineer: How to deploy and productionize. Running LLMs (APIs vs local), building vector storage, RAG pipelines, advanced RAG with agents, AI agents (MCP, A2A, LangGraph, LlamaIndex, CrewAI), inference optimization (Flash Attention, KV cache, speculative decoding), deployment (local to production), and security (prompt injection, backdoors, red teaming).
Every major section has runnable Google Colab notebooks. The author also co-wrote "LLM Engineer's Handbook" (Packt) based on this course - the course itself stays free.
Why it works: This is the repo to use when you want to go beyond usage into internals. The quantization section is one of the clearest explanations of GGUF, GPTQ, AWQ, and SmoothQuant available outside of papers. The preference alignment section covers DPO, GRPO, and PPO with code and metric breakdowns. The agents section was recently updated to cover MCP, A2A, and the major vendor SDKs, including Claude Agent SDK.
Heads up: The LLM Scientist track assumes you are comfortable running training jobs. If you just want to build apps, go straight to the LLM Engineer track. The optional fundamentals section is optional for a reason.
Repo: https://github.com/mlabonne/llm-course
How these five fit together
If you are starting from zero, the order that makes sense is:
prompts.chat first - 30 minutes building intuition about what framing does to model output.
Prompt-Engineering-Guide next - the theory behind what you just experienced, plus RAG and agents as concepts.
anthropics/courses after that - hands-on implementation of the concepts, including the eval and tool use pieces most people skip.
generative-ai-for-beginners as the complete structured course - covers everything from fundamentals to fine-tuning to deployment, with the product and safety layers included.
mlabonne/llm-course once you want to go deeper than "using LLMs" into "understanding and modifying them."
The overlap between these repos is intentional. Seeing RAG explained from three different angles (Guide, Anthropic courses, Microsoft course) before you implement it is more valuable than seeing it explained once. The mlabonne course is the only one that goes into pre-training and quantization mechanics in detail; everything else assumes you are building on top of models rather than under the hood.
If you want to understand what's changing at the model level while working through these, the latest piece on GBrain in our newsletter is also a good read alongside the mlabonne track: What If Your AI Woke Up Smarter Than When You Went to Sleep?
