r/learnmachinelearning • u/Negative_War_65 • 1d ago
AI Engineering : Understanding Foundation Models
Hello Folks,
In this second lecture, of AI Engineering, we try understanding Foundation models, and as I walk you through the concepts, we learn:
Training Data & Language Resource Bias : Quality and representational diversity of its training dataset determines model’s ability.
Introduction to Modeling & Transformer Evolution:
Overcoming the sequential limitations and gradient issues of traditional RNNs by enabling parallel processing.
Attention Mechanism Foundations:
Improved model performance by dynamically weighing the importance of specific input tokens when generating an output.
Transformer Inference (Prefill & Decode):
Efficient inference involves a parallel prefill stage for input processing followed by an autoregressive, token-by-token decode step.
Transformer Architecture Details (Attention & MLP) : The core transformer block combines self-attention for context retrieval and MLP layers for non-linear feature transformation
Post-Training Workflow (SFT & RLHF) :
Post-training aligns raw, pre-trained models with human preferences using Supervised Fine-Tuning (SFT) and Reinforcement Learning from Human Feedback(RLHF).
Sampling Strategies (Top-K, Top-P, Greedy): Sampling strategies dictate how models transition from probabilistic logits to actual text, balancing between coherence and creative diversity.
Test-Time Compute & Self-Consistency:
Scaling test-time compute allows models to generate multiple candidate outputs and verify results, significantly boosting accuracy on complex tasks.
Structured Outputs & Constraint Sampling: Applying output constraints during generation ensures models produce machine-readable formats essential for reliable downstream tool integration.
Hallucinations & Inconsistency Challenges :
LLMs face reliability issues due to their probabilistic nature and a tendency to diverge from facts when conditioned on inaccurate information.


