Frontend
MongoDB Aggregation Pipelines: Stage Order Is the Win
Yaseen Khatib Dev.to (EN Zone)
3 views
[ EXECUTIVE TEARDOWN // TL;DR ]
Put $match first and on indexed fields — every document eliminated early is one the rest of the pipeline never touches.
$lookup is a join: run it late (after match/limit) and index the foreign field, or it degrades to repeated scans.
Profile with explain('executionStats') and reorder by the real numbers, not intuition.
Design pipelines bottom-up: find the smallest set, then ensure every stage operates only on it.
On a clinical workflow API, an aggregation pipeline that had always felt instant started to crawl — same query, same result, just more data flowing through it before anything got filtered out. That is the whole game with MongoDB aggregation: identical logical output can vary by orders of magnitude depending on stage order and index usage, and tightening these pipelines is where the real speedup lives — no bigger instance required.
Filter early, filter on indexes
The single most important rule: put $match first, and match on indexed fields. Every document you eliminate at stage one is a document the rest of the pipeline never touches. A pipeline that sorts or joins before filtering is doing expensive work on rows it is about to throw away — the database equivalent of cleaning a house before deciding to demolish it.
pipeline.ts
// $match first, on an indexed field — shrink the set early
const rows = await Visits.aggregate([
{ $match: { clinicId, status: "active" } }, // indexed, runs first
{ $sort: { scheduledAt: -1 } }, // on the small set
{ $lookup: { from: "patients", localField: "patientId",
foreignField: "_id", as: "patient" } },
{ $limit: 50 },
]);
$lookup is a join — treat it like one
$lookup is the most expensive stage in most pipelines because it is a join, and joins multiply work. Run it as late as possible — after matching and limiting — so you join against dozens of documents, not thousands. And ensure the foreign field is indexed, or each lookup degrades into a collection scan repeated once per input document.
Profile with explain
Guesswork has no place here. Run the pipeline through explain("executionStats") and read what actually happened: which stages used an index, how many documents each examined versus returned, and where the time went. The speedup came from reading those numbers and reordering stages — not from intuition about what "should" be fast.
An aggregation pipeline is read top to bottom but should be designed bottom-up: decide the smallest set you can get away with, then make sure every stage operates only on it.
The production system is the Hospital-API; for the vector side of querying, see Vector Embeddings in Production . Pipelines like these are the unglamorous backend work behind the five products I shipped solo this year — the judgment I bring is knowing the fastest query is usually a reordering problem, not a bigger box.
~/keep-reading
7 min readSerialization Adapters: How I Cut Payloads by 94%Rich UI objects make terrible database records. A Serialization Adapter I built for IntegrateX split render model from transport record and cut payloads by 94%.
6 min readCutting a Payload 94% With Custom Serialization PatternsI cut a React Flow agent-graph payload 94% without losing a node — not with gzip, but by shaping a custom serialization format around the data: schema, not prose.
8 min readWebSocket Telemetry at Scale: When One Process Isn't EnoughA single WebSocket server is a weekend project; streaming telemetry to thousands across instances broke for me on streamerOS — Redis pub/sub, rooms, coalescing.
YK
Yaseen Khatib · MERN + AI Architect
Ships autonomous AI products solo — five in the last twelve months. More about Yaseen →
Need an engineer who can build this?
I'm Yaseen Khatib — a Senior Full-Stack AI Engineer (MERN + TypeScript) who ships production AI systems solo. Open to senior and lead roles, remote or on-site.
Get in touch →See what I've shipped
Originally published at yaseenkhatib.streamerosai.com/blog/optimizing-mongodb-aggregation/.
Read original: https://dev.to/yaseenyk04/mongodb-aggregation-pipelines-stage-order-is-the-win-3cfk
← Previous
WebSocket Telemetry at Scale: When One Process Isn't Enough
Next →
I Built an AI Money Agent That's Structurally Incapable of Touching the Money
Related
[Showoff Saturday] A new portfolio page!
Frontend
1
Reddit r/webdev
How I Built a Fast, Clean Wiki & Database for Steal an Egg
Frontend
4
Dev.to (EN Zone)
Got tired of writing READMEs, so I built a tool that does it for me
Frontend
4
Reddit r/webdev
How to catch the frozen tabs and jank your monitoring misses
Frontend
5
Dev.to (EN Zone)
Comments0
No comments yet — be the first