Interactive walkthroughs of real AI infrastructure stacks: the actual repos and services involved, how data moves between them, and where storage and memory sit at each stage. Updated as new systems are worth covering, not on a weekly schedule.
Data curation through training, parallelism, post-training, evaluation, export, and serving, across the modular NeMo Framework repos (Curator, AutoModel, Megatron-Bridge, RL, Evaluator, Export-Deploy, Speech), plus how each stage runs on Kubernetes.