Skip to content

A home for your AI & ML interview prep

You’ve collected the resources. Now turn them into something you can build, explain, and take into an interview.

A little direction. Then a lot of doing.

First, a small win

You probably know more than you think.
Put one decision to the test.

What’s getting in the way?

On the desk: GenAI / RAG. The docs have the answer. Your bot doesn’t.

A little hands-on practiceGenAI / RAG

The docs have the answer.
Your bot doesn’t.

Your support bot gives the wrong policy, even though the correct one is in your documents.

Where would you look first?

Keep practising RAG

What you can work on

The foundations.
And everything you build on them.

28 topics, connected to questions, coding labs, and design problems. Follow a path or start with a gap.

  1. 01

    The mathematical foundations

    Linear algebra, probability, optimization, information theory, and numerical computing.

  2. 02

    How models learn

    ML fundamentals, classical models, evaluation, and the reasoning behind your choices.

  3. 03

    From neural networks to attention

    Deep learning, transformers, generative models, and multimodal architectures.

  4. 04

    Build with language models

    RAG, embeddings, fine-tuning, agents, alignment, and evaluation.

  5. 05

    Make it work at scale

    Model serving, distributed training, quantization, and the trade-offs of production.

Explore the topic library

Now, give it a direction

Your next interview.
One step at a time.

  1. Plan
  2. Design
  3. Code
  4. Talk it through

01 Make a plan

An interview on the calendar.
A place to begin.

Your experience, the role, and the time you have. Turn them into a prep plan with a finish line.

Start where you are

Bring your resume and the job description.
We’ll help you find what to work on next.

Make it my plan

02 Think in systems

Make the pieces make sense.

Start with a few building blocks. Use a hint, complete the request path, and see how your design holds up.

The system design canvasLLM inference Demo
Hint The gateway accepts a request. What decides which model should handle it?
Place the missing piece. Connect the reasoning.
✓ Routing connectedA clear request path, with a fallback.92/100

From a missing connection to the trade-offs behind a complete architecture.

Open the LLM gateway design lab

03 Build the understanding

From “I get it” to “it runs.”

Write multi-head attention with a little help along the way. Follow the dimensions, run the checks, and see the shapes line up.

attention.pyPyTorch Demo
AI hint Two examples, four tokens, 64 features, 8 heads. Split the heads, then bring them back together.

# X: (2, 4, 64), requires_grad=True · W*: (64, 64)

Q, K, V = [X @ W for W in (Wq, Wk, Wv)]
Q, K, V = [t.reshape(2, 4, 8, 8).transpose(1, 2)
           for t in (Q, K, V)]
scores = Q @ K.transpose(-2, -1) / (8 ** 0.5)
allowed = torch.ones(4, 4, dtype=torch.bool).tril()
weights = scores.masked_fill(~allowed, -torch.inf)
heads = torch.softmax(weights, dim=-1) @ V
out = heads.transpose(1, 2).reshape(2, 4, 64) @ Wo
out.square().mean().backward()
Output▷ Run checksRunning checks…✓ Shapes and gradients checked
Q, K, V (2, 8, 4, 8)Attention (2, 8, 4, 4)Output (2, 4, 64) ✓

AI assistance when you need a nudge. Real coding labs when you’re ready to try.

Explore the attention coding lab

04 Say it out loud

Now, talk someone through it.

Meet Alex, your AI interviewer. Explain what you built, answer the follow-up, and leave with something specific to improve.

A conversation with AlexMock interview Demo
  1. Alex: Walk me through the shapes in your multi-head attention implementation. You: I split 64 features across 8 heads. Each head still sees all 4 tokens, with 8 features per token.
  2. Alex: Why divide the attention scores by the square root of the head dimension? You: It controls the scale of the dot products as the dimension grows, so softmax is less likely to saturate.
  3. Alex: And how do you get back to the original output shape? You: Concatenate the 8 heads into 64 features, then apply the output projection. The result is batch × 4 × 64.
Feedback

Clear on shapes and scaling. Next, explain where you would apply a causal mask.

A conversation that gets beyond the first answer.

Explore mock interviews

You don’t have to do it all today

Just take
the next step.

One question. One working idea.
A little more ready than yesterday.

Build my prep plan