AI Engineer Path

Week 38

Project polish

Turn eight projects into evidence: diagrams, tradeoffs, eval results, cost analysis, demos.

Why this week matters

For each project: an architecture diagram, stated tradeoffs, eval results and cost analysis. The last two are what separate you from most applicants: almost nobody includes them.

Done when

Both flagships have a README a hiring manager can scan in two minutes, with numbers, a diagram and a live demo link.

Concepts

5 lessons · tick each one once you could explain it

In plain words

Your README is often the first and only thing a hiring manager reads, usually for about two minutes. Open with what the project does and the proof it works (a demo, a GIF, the key numbers), then explain how it's built and the choices you made, and finish with honest limitations.

A shop window: the best product, its price tag and a 'try me' sign up front; the stockroom inventory goes at the back.

In detail

Structure: one-sentence summary; a GIF or screenshot and live demo link; results (eval table, cost per request, latency); architecture diagram; key design decisions with tradeoffs; how to run it; limitations and next steps.

Write for a reader with two minutes. Numbers and visuals near the top; implementation details further down.

Honesty is a feature: 'faithfulness 0.91 on 150 questions; fails on multi-hop questions spanning several documents' is more credible than 'highly accurate'.

Worked example

Rewriting the top of a README

  1. Before: 'A RAG system for HR policies built with Python and LangChain.' It tells the reader nothing about quality.
  2. After: 'Answers HR policy questions with citations. Faithfulness 0.91 on 150 questions, p95 latency 2.3 s, 0.004 USD per query.'
  3. Directly under that: a live demo link and a 10-second GIF.
  4. Then the results table, an architecture diagram and 'Design decisions' (for example: pgvector over Qdrant, and why).
  5. Last: how to run it, plus 'Limitations: struggles with questions spanning several documents.'

What it does and the proof first; how and why next; honest limits last.

Common mistakes

  • Opening with installation steps or the tech stack instead of what the project does and how well.
  • Vague claims ('highly accurate') with no numbers to back them.

Check yourself

Why include limitations in a portfolio README?Show answer

They show you measured the system and understand it. Specific honesty is more credible than vague praise, and reviewers trust the rest more.

What goes in a 'Design decisions' section?Show answer

Three to five key choices, each with the rejected alternative and the reason, which shows engineering judgement.

Going deeper

Add a 'Design decisions' section with 3–5 choices and their rejected alternatives ('pgvector over Qdrant because…'). Reviewers look for judgement, and this is where it shows.

Best resources for this lesson

In plain words

An architecture diagram shows, on one page, how a request travels through your system and where data is stored. Draw the live request path and the offline data-preparation path separately, and label the databases, outside services and where monitoring connects. You'll redraw it on a whiteboard in interviews.

A metro map: not every street, just the lines, the stations and where you change trains.

In detail

Draw the main request flow left to right (client → API → retrieval → reranker → LLM → response) and the offline flow separately (ingest → chunk → embed → index). Label stores, external APIs, queues and where tracing and evals plug in.

Tools: Mermaid in the README (renders on GitHub), Excalidraw or draw.io.

A diagram prepares you for the system design interview: you'll redraw it on a whiteboard.

Worked example

A RAG system on one page

  1. Online path, left to right: User → FastAPI /ask → hybrid retrieval (BM25 + pgvector) → cross-encoder rerank → LLM with citations → response.
  2. Offline path, drawn separately: documents → chunk → embed → index into pgvector.
  3. Dotted lines from the API and LLM to Langfuse show where traces go.
  4. Write it in Mermaid inside the README, so GitHub renders it and it's versioned with the code.
  5. Practise drawing it from memory in under two minutes.

One container-level diagram with online and offline paths, stores and observability is usually exactly right.

text
flowchart LR
  U[User] --> API[FastAPI /ask]
  API --> R[Hybrid retrieval<br/>BM25 + pgvector]
  R --> RR[Cross-encoder rerank]
  RR --> LLM[LLM w/ citations]
  LLM --> API
  API -. traces .-> LF[(Langfuse)]

Common mistakes

  • Cramming every class and function into the diagram until nothing is readable.
  • Leaving the offline ingestion path out, which hides half the system.

Check yourself

Why use Mermaid rather than an image file?Show answer

It's text, so it lives in the repository, shows up in diffs, stays in sync with the code and renders natively on GitHub.

What level of the C4 model suits a portfolio README?Show answer

The container level: services, databases and external systems, without code-level detail.

Going deeper

The C4 model gives levels of zoom (context, containers, components, code); one container-level diagram is usually right for a portfolio README. Mermaid renders natively on GitHub, so diagrams stay versioned with the code.

In plain words

Report evaluation results as a small table: one row per version of your system (starting with the baseline) and one column per metric that matters, plus cost and latency. Say how big the test set was, how it was labelled and how you checked the judge. That one table shows you can improve a system and prove it.

A before-and-after chart in a fitness diary, with the scales checked for accuracy.

In detail

Rows are versions (baseline, +hybrid, +rerank, final); columns are the metrics that matter (retrieval recall@5, faithfulness, answer correctness, abstention accuracy) plus cost and latency.

State the eval set's size and composition, how it was labelled, and how the judge was validated.

This one table shows you can improve a system and prove it, which is the core LLMOps skill.

Worked example

A results table that convinces

  1. Rows: naive RAG, plus hybrid search, plus reranking.
  2. Recall@5: 0.71 → 0.83 → 0.89. Faithfulness: 0.82 → 0.85 → 0.91. Correctness: 0.64 → 0.71 → 0.78.
  3. Plus p95 latency and cost per 1,000 queries in the same table, so the trade-offs are visible.
  4. Under it: 150 questions (120 answerable, 30 not), labelled by you; judge agreement with your labels 88%.
  5. Add the uncertainty: at 0.78 on 150 questions, the standard error is about ±3.4 points.

Baseline row first, versions as rows, metrics plus cost and latency as columns, and the eval set described honestly.

text
| Version          | Recall@5 | Faithful | Correct | p95 latency | Cost/1k q |
|------------------|---------:|---------:|--------:|------------:|----------:|
| Naive RAG        |   0.71   |   0.82   |  0.64   |    2.1 s    |   USD 1.90 |
| + hybrid search  |   0.83   |   0.85   |  0.71   |    2.3 s    |   USD 1.95 |
| + rerank (final) |   0.83   |   0.92   |  0.79   |    2.6 s    |   USD 1.40 |

Common mistakes

  • Reporting only the final version, which hides how much each change helped.
  • Numbers without the eval-set size or labelling method.

Check yourself

Why put the baseline row first?Show answer

Every improvement is measured relative to it; without a baseline, a number like 0.78 has no meaning.

Why mention the eval-set size next to the numbers?Show answer

It shows how much noise to expect: on 150 questions a 2-point change may be random, and readers can judge that.

Going deeper

Include confidence intervals or at least the eval-set size next to every number, and one or two qualitative examples of failures. It signals statistical literacy and honesty.

Best resources for this lesson

In plain words

Show what your system costs per request and per month at realistic volumes, using the token counts you logged. Then list the ways you'd cut the cost and how much each would save. Few portfolios do this, and it maps directly onto a common interview question.

A business plan's cost sheet: what each sale costs to make, and how you'd make it cheaper.

In detail

Compute from logged tokens: average input and output tokens per request × prices, plus embedding, reranking and infrastructure costs. Project to a monthly volume (e.g. 10k and 1M requests per month).

Then show levers with estimated savings: prompt caching, routing easy queries to a smaller model, fewer but better chunks after reranking, semantic caching, batch APIs for offline work.

This section is rare in portfolios and maps directly to the 'cut the bill by 60%' interview question.

Worked example

Per-request and monthly cost

  1. Logged averages: 2,400 input and 300 output tokens. At 3 and 15 USD per million: 0.0072 + 0.0045 = 0.0117 USD.
  2. Add embeddings and reranking, about 0.0008, for a total of about 0.0125 USD per request.
  3. Monthly: 10,000 requests ≈ 125 USD; 1 million requests ≈ 12,500 USD.
  4. Lever: caching the 1,500-token system prompt at about a tenth of the price saves about 0.004 per request, roughly 32%.
  5. Cost per successful task: with 92% success, 0.0125 / 0.92 ≈ 0.0136 USD.

Cost per request from logged tokens, monthly projections and ranked levers with estimated savings.

Common mistakes

  • Estimating from guessed token counts instead of logged ones.
  • Ignoring retries and failures, which also cost money.

Check yourself

Why report cost per successful task rather than per request?Show answer

Failed and retried requests also cost money, so cost per success reflects the true unit economics.

Name three cost levers in order of increasing risk.Show answer

Prompt caching and context trimming (low risk), routing easy traffic to a smaller model (needs evals), then distillation or self-hosting (high effort).

Going deeper

Show unit economics: cost per successful task (not per request), since retries and failures cost money too. Add a sensitivity line: 'if traffic 10×, cost becomes… with caching becomes…'.

Best resources for this lesson

Where this comes back

In plain words

Your smaller projects need the same care as the flagships, just shorter: a summary, the key metric, a demo or screenshot and one honest limitation. Use the same README structure everywhere, pin your six best repositories, and hide or archive half-finished ones that dilute the impression.

Tidying the whole house before guests arrive, not just the living room.

In detail

Document-understanding project: field-level accuracy table, cost per page, a gallery of tricky documents. Fine-tuning project: the before/after table against strong prompting, and your verdict.

Pin the best six repositories. Archive or make private anything half-finished that dilutes the signal.

Consistency matters: the same README skeleton across projects makes the portfolio feel like the work of one careful engineer.

Worked example

A portfolio clean-up afternoon

  1. Document-extraction project: add a field-accuracy table (93%), the cost per page (0.004 USD) and a gallery of tricky documents.
  2. Fine-tuning project: add the before-and-after table against strong prompting, and your verdict.
  3. Apply the same README skeleton to both: summary, demo, results, architecture, decisions, how to run, limitations.
  4. Pin six repositories: two flagships, GPT from scratch, extraction, fine-tune and the MCP server.
  5. Archive three half-finished repositories so they don't distract.

Every visible project gets a metric and an honest limit, in a consistent format; unfinished work gets archived.

Common mistakes

  • Leaving abandoned experiments pinned or public next to your best work.
  • Inconsistent READMEs that make the portfolio feel scattered.

Check yourself

Why use one README template across projects?Show answer

Reviewers find information in the same place every time, and the portfolio reads as the work of one careful engineer.

What's the minimum each supporting project needs?Show answer

A one-line summary, its key metric, a demo or screenshot, how to run it and one honest limitation.

Going deeper

A consistent template across repos (summary, demo, results, architecture, decisions, run, limitations) makes the portfolio read as one careful engineer's work.

Best resources for this lesson

Practice

Hands-on work that makes the lessons stick. Warm-ups take minutes; stretch goals are optional.

  1. Core

    Two-minute README test

    Give a friend two minutes with each flagship README and ask what it does, how well, and at what cost. Fix whatever they couldn't answer.

  2. Warm-up

    Mermaid diagrams

    Add a container-level Mermaid diagram to each flagship README, showing online and offline flows.

  3. Stretch

    Cost sensitivity

    Add a cost table at 10k, 100k and 1M requests/month with and without caching and routing.

This week, day by day

Dates follow your pace from Settings. Open the notebook icon to log hours and notes.

  1. Day 260Monday7 Jun2 h planned

    Flagship 1: README, architecture diagram, eval results

  2. Day 261Tuesday8 Jun2 h planned

    Flagship 1: cost analysis and live demo link

  3. Day 262Wednesday9 Jun2 h planned

    Flagship 2: README, diagram, sample traces

  4. Day 263Thursday10 Jun2 h planned

    Flagship 2: cost analysis and live demo link

  5. Day 264Friday11 Jun2 h planned

    Polish the document-understanding project

  6. Day 265Saturday12 Jun3 h planned

    Polish the fine-tuning project

  7. Day 266Sunday13 JunReview

    Review the week, finish anything unfinished, rest

Read and use

Interview prep

Questions this week's material gets asked as. Answer out loud first, then open the outline.

Walk me through your best project.
  • Problem and why it matters
  • Architecture and key decisions with alternatives
  • Eval results, cost, limitations; what you'd do next

Check yourself

Five questions. The done-when test above is the real bar; this is a quick self-check.

  1. 1.What should appear near the top of a project README?

  2. 2.Which two portfolio elements do almost no applicants include?

  3. 3.An eval results table should compare versions on…

  4. 4.Stating a known limitation in your README…

  5. 5.A cost analysis should include…