"Design an API for this" is one of the few interview exercises where candidates lose points for answering quickly.
The prompt is an invitation to list endpoints, and a list of endpoints is what most people produce: POST /shipments, GET /shipments, GET /shipments/{id}. All correct, all obvious, and none of it is the part anyone gets wrong in production. The decisions that actually cost teams months are the four below, and none of them appear in the obvious list.
For the fundamentals this round assumes — verbs, status codes, idempotency, versioning, REST against gRPC — the question set is here: backend developer interview questions. This page is the design exercise itself.
The prompt
"We're exposing a public API for our shipping product. Customers create shipments, print labels, and track deliveries. Design it."
Ask one question first, and make it this one:
"Who's calling it? A public API for customers' own engineers is a different design from an internal one — the difference is that I can never change it, so the things I'd defer on an internal API I have to decide now. I'll assume public and long-lived."
That's the whole framing, and it justifies everything that follows. A candidate who designs a public API with internal-API care has misunderstood the cost structure.
The obvious part, quickly
"The resources are shipments, labels and tracking events. Shipments are the aggregate, labels belong to a shipment, tracking events are append-only and read-only to the customer.
So: create and read shipments, request a label for one, list tracking events for one. I'd rather spend the time on four things that aren't in that list."
Saying this in twenty seconds and moving on is itself a signal — you've shown you know the obvious answer and that you know it isn't the interesting part.
Decision one: how lists are paginated
"GET /shipments needs pagination, and the choice between offset and cursor is the first real decision.
Offset — ?page=3&per_page=50 — is easy to build and easy to document, and it's wrong for this. Shipments are created constantly. If a customer pages through while new ones arrive, rows shift between pages: they'll see duplicates and miss records, and they'll never report it as a bug because it looks like their own data.
So cursor: ?limit=50&after=<opaque>. The cursor encodes the sort key and the ID for tie-breaking. Opaque on purpose — the moment customers parse it, I can't change the sort implementation.
The cost is real and worth naming: no jumping to page 40, and no total count without a second query. For a shipping API I'd take that trade. For something where a UI needs page numbers, I might not."
Naming the cost of the choice you made is the whole exercise. Both pagination styles are defensible; a candidate who picks cursor and can't say what it gives up has repeated a recommendation rather than made a decision. The interviewer is not checking which one you chose — they're checking whether you know what breaks either way, because that's what tells them how you'll behave on the next decision, which they won't be there for.
Decision two: the error shape
"I'd fix the error format before writing a second endpoint, because it's the thing you can never change afterwards and the thing every client wires into.
Every error returns the same body: a machine-readable code, a human message, and a field list for validation failures. Something like code: \"address_invalid\", message, and errors as an array of field plus reason.
The important part is that the code is stable and the message is not. Clients switch on the code; the message is for the developer reading the log, and I want to be free to improve it. If I don't say that in the documentation, someone will parse the message string and I'll have broken them the first time I fix a typo."
"And I'd include a request ID in every response, success or failure. When a customer emails support with 'it didn't work', that ID is the difference between a five-minute answer and an afternoon."
Decision three: bulk operations that half-fail
"Customers will want to create shipments in batches, and this is where most APIs go wrong.
The temptation is POST /shipments/bulk returning 200 or 400. But a batch of fifty where three are invalid is the normal case, not the edge case, and both answers are wrong — rejecting all fifty for three bad rows is infuriating, and returning 200 hides the failures.
What I'd do is return 207 with a per-item result: index, status, and either the created resource or the error object from the format above. The client iterates and retries the three.
The subtlety is idempotency: if the client retries the whole batch, the forty-seven that succeeded must not be created twice. So each item carries its own idempotency key, not one for the batch."
That last paragraph is the one that shows experience. Batch idempotency at the wrong granularity is a bug that ships and then duplicates customer orders.
Decision four: work that takes longer than a request
"Label generation calls a carrier, which can take ten seconds or fail. I wouldn't put that behind a synchronous POST /labels that blocks.
So: POST /shipments/{id}/labels returns 202 with a label resource in state pending and a URL to poll. The label becomes ready with a PDF URL, or failed with a code from the carrier.
And I'd add webhooks, because polling is a bad experience at scale and every integration eventually asks for them. Webhooks are their own design problem — retries, signing, ordering — and if the interviewer wants to go there it's a better conversation than more endpoints. The one thing I'd commit to now is that the webhook payload contains the resource ID and the event type, not the full resource, so I'm not locked into a payload shape I have to version separately."
The endpoint list
“POST /shipments, GET /shipments, GET /shipments/id, POST /labels, GET /tracking. Standard REST, JSON, status codes, JWT auth.”
Complete, correct and indistinguishable from every other candidate's answer. Nothing in it could be wrong, which means nothing in it demonstrates judgement — and the interviewer still has twenty-five minutes to fill.
Four decisions with costs
“Cursor pagination, and here's what it costs. A frozen error shape. 207 with per-item idempotency for batches. 202 and a poll for label generation.”
Each one is a fork with a stated trade-off, and each is something that genuinely breaks in production. The conversation now has somewhere to go, and every direction is ground you chose.
The pushbacks to expect
"Isn't 207 unusual? Most APIs don't use it." — hold the position and concede the fair part: "It is unusual, and if the team preferred a 200 with the same per-item body I wouldn't argue hard — the status code matters less than the fact that partial success is representable. What I'd refuse is an all-or-nothing batch, because that pushes the retry logic onto every customer."
"What about rate limiting?" — they're checking whether you think past the happy path. Headers on every response rather than only on the rejection, so a well-behaved client can slow down before it gets a 429. And a limit scoped per customer rather than per IP, because customers behind one gateway shouldn't share a bucket.
"How would you let customers filter the list?" — the trap is inventing a query language. "A small fixed set of filters — status, date range, destination country — and nothing composable. The moment I accept arbitrary field filtering I've promised to keep every field indexable forever."
Preparing for this round
Take an API you've actually built and write down the four decisions above as they were made. Most people find that at least one of them was never decided — it was inherited from a framework default, which is a perfectly good answer as long as you know it.
Then practise saying the cost out loud. The instinct under interview pressure is to present choices as obviously correct, and the correction is small: after each decision, one sentence beginning "what that costs is…".
Practical target: open the documentation of an API you actually call and find the decision that was clearly made under time pressure — the endpoint that returns a different shape from its siblings, the list without pagination, the error that's a bare string. Then work out what it would take to change it now. That's the clearest available lesson in why this round exists, and it takes about fifteen minutes.




