Who it’s for
Learners looking for a hands-on introduction to how language models work, from the first concepts to a working project.
Fine-tune a real model. Ship a demo in 6 weeks.
For beginners with no ML background. Use guided notebooks to fine-tune Needle, a 26-million-parameter model, and build a web demo you can share.
FIND YOUR FIT
Learners looking for a hands-on introduction to how language models work, from the first concepts to a working project.
No prior machine learning background is required. The course introduces the concepts and uses guided notebooks.
MAKE IT REAL
Choose a focused task, prepare examples, and fine-tune Needle, a 26-million-parameter open-source model. Wrap it in a simple web demo and share it on demo day.
THE TAKEAWAY
Explain how text, tokens and attention fit together in a language model.
Explore a real model and its data in Google Colab.
Create a focused dataset and fine-tune Needle for your chosen task.
Build a clickable demo and present what you made.
MEET YOUR INSTRUCTOR
HOW THE PROGRAM WORKS
Explore each week to see the sessions, topics, and teaching dates. All times are in IST.
The big idea behind every language model: predicting the next word, over and over, at huge scale. What Needle, a real 26-million-parameter model, tells us about how small "useful" can actually be.
Setting up Google Colab, running your first notebook, and loading Needle to watch it pick a tool and fill in the arguments.
Tokenization explained simply: breaking text into pieces a model can learn from, and why vocabulary size matters.
What training data actually looks like, using Needle's own tool-calling examples as a real, inspectable dataset.
The one idea behind every modern model: letting each word "look at" the others for context. Needle makes an unusually good teaching example, since its architecture drops the usual extra layers and keeps only attention. There's less to explain.
Assembling a small attention block in code, guided line by line. You don't need to derive it, just see it work and inspect it.
Loss, learning rate, and epochs, explained as "how a model finds out it's wrong and gets a little better every time." Plain language, with a training run you can actually watch.
Needle was pretrained on 200 billion tokens over about 27 hours, then specialized in roughly 45 minutes of fine-tuning. Why the second number is the one that matters for this course, and for you.
Choosing something specific and useful for your model to do, and building a small, focused dataset of examples for it.
Running Needle's own fine-tuning playground on your data, and watching your general-purpose model become specific to your task.
Wrapping your fine-tuned model in a simple, clickable web demo anyone can try. The same instinct behind Needle's own phone-ready design.
Present what your model does, how you built it, and what you'd try next, to classmates, faculty, or anyone willing to test it live.
BEFORE YOU JOIN
Session recordings are included with the program.
Sign in and open your dashboard to access your enrolled programs.
Read the refund policy before enrolling. For questions about your situation, contact hello@gradlabs.in.
YOUR NEXT STEP
Review the current fee and reserve your place, or ask us a question before you decide.