Limited time: AI code review, hints, mock interviews, whiteboard analysis, and all Pro features are unlocked. Enroll
โฑ๏ธ 16 min read

LLM APIs and SDKs - Complete Deep Dive

Prerequisites: How LLMs Actually Work, Python for AI Engineers Used in: Tokens and Cost Math, Structured Outputs, Tool Calling, LLM Observability Build it: Lesson 2 - The Model Boundary implements this as runnable, tested code you can execute offline.


What is the LLM API surface?

Underneath every AI product is a single HTTP endpoint that takes a list of messages and returns one more message. That is genuinely the whole abstraction. The major providers โ€” OpenAI, Anthropic, Google, and the open-weight hosts serving Llama, Qwen, and Mistral โ€” arrived at near-identical request shapes, so learning one transfers almost completely to the others. Field names and a few flags differ; the model of interaction does not.

The property that surprises people is that the endpoint is stateless. It keeps no conversation for you. There is no session, no server-side thread you append to. Every call is a complete restatement of the entire conversation so far, and the model is effectively a pure function from the message list to the next message.

Real-world analogy: A brilliant consultant with no memory between meetings. Every time you walk in, you hand over the complete case file โ€” the standing instructions, everything both of you have said, and any documents you want considered โ€” and they respond based only on what is in front of them. If you forget to include a page, it does not exist. You pay by the page you hand over and by the page they write back.

Conversation state is therefore your problem, held in your database and reassembled on every request. This is also why cost grows quadratically across a long chat if you never trim: turn 20 resends turns 1 through 19.


The Converged Request Shape

{
  "model": "your-chosen-model",
  "messages": [
    { "role": "system",    "content": "You are a support triage assistant. Reply in JSON." },
    { "role": "user",      "content": "My card was charged twice." },
    { "role": "assistant", "content": "I can help. What is the order number?" },
    { "role": "user",      "content": "Order 4471." }
  ],
  "temperature": 0,
  "max_output_tokens": 512,
  "stream": false
}
Role Who writes it What it is for
system You, never the user Standing instructions, persona, output contract, policy. Sent on every call.
user The end user, or you on their behalf The turn to respond to, plus any retrieved context you injected
assistant The model, replayed by you Prior model turns, so the model sees its own commitments
tool Your code The result of executing a tool the model asked for, keyed to its call id

Two practical rules fall out of the role split. Never put user text in the system message โ€” concatenating them is how instruction-override attacks land, covered in AI Security. And never trust the client to send the system prompt or the history; assemble both server-side from your own store, exactly as you would not trust a browser to send the price of an item.


Parameters That Matter

Parameter What it controls Sane default
temperature Randomness of sampling. Low is focused and repeatable, high is varied 0 for extraction, classification, and tool selection. Raise only for genuinely creative copy
top_p Nucleus sampling cutoff over the probability mass Leave at 1. Tune temperature or top_p, never both at once
max_output_tokens Hard ceiling on generated length Always set it explicitly. It is your per-request cost ceiling and your runaway-generation guard
stop sequences Strings that end generation early Useful when you delimit output yourself. Unnecessary with schema-constrained output
seed Reproducibility hint where a provider offers it Set it in evals. Treat it as best-effort, not a determinism guarantee
n Number of independent samples Leave at 1. It multiplies output cost linearly, so use it only for deliberate self-consistency work

The default worth internalising is temperature: 0 for anything a program consumes. Engineers routinely leave a provider default of around 1 in place for a JSON extraction task and then spend a day debugging โ€œflaky parsingโ€ that was sampling noise the whole time.

Note that temperature: 0 gives you more repeatable output, not guaranteed identical output. Batching, mixed-precision kernels, and fleet heterogeneity all introduce variation, so write tests that assert on meaning rather than on byte equality.


Streaming vs Buffered

A buffered call returns one response after the full generation completes. A streaming call returns incremental deltas over a long-lived connection as tokens are produced, usually as server-sent events.

flowchart LR
    U[User] --> API[Your API]
    API --> ASM[Assemble messages from store]
    ASM --> P[Provider endpoint]
    P --> T1[First delta after TTFT]
    T1 --> T2[Deltas stream in]
    T2 --> DONE[Final chunk carries finish reason and usage]
    T1 --> U
    DONE --> LOG[Persist full text and token usage]

    classDef client fill:#f97316,stroke:#c2410c,color:#fff
    classDef edge fill:#6cf,stroke:#333,color:#000
    classDef service fill:#10b981,stroke:#065f46,color:#fff
    classDef async fill:#b4f,stroke:#333,color:#000
    classDef data fill:#fbbf24,stroke:#92400e,color:#000

    class U client
    class API edge
    class ASM,P service
    class T1,T2,DONE async
    class LOG data
ย  Buffered Streaming
Perceived latency Full generation time Time to first token, typically a small fraction of it
Client complexity One response object Incremental consumer, reassembly, disconnect handling
Validation Validate the whole payload before acting Cannot validate until complete, so risky for tool calls and strict JSON
Failure mid-response Clean error, nothing shown Partial output already on screen
Cancellation Wasted spend Can abort and stop paying for further tokens

Stream when a human reads the output as it appears. Stay buffered when a program consumes it โ€” half a JSON object is not useful, and you would only buffer it yourself anyway. The usage numbers arrive in the final chunk, so a stream you abandon early leaves you without token counts unless you estimate them; log the partial either way. The system-design view of streaming at scale, including fan-out and connection management, is in Design ChatGPT and WebSockets.


Tool Calling at the API Level

At the request surface, tool calling is small: you pass a list of function schemas alongside your messages, and the model may respond with a structured request to call one rather than with prose.

flowchart LR
    R[Request with messages and tool schemas] --> M[Model]
    M --> D[Finish reason is tool call]
    D --> EX[Your code executes the tool]
    EX --> APP[Append assistant tool call and tool result]
    APP --> M2[Second model call with full history]
    M2 --> ANS[Final assistant message]

    classDef client fill:#f97316,stroke:#c2410c,color:#fff
    classDef service fill:#10b981,stroke:#065f46,color:#fff
    classDef async fill:#b4f,stroke:#333,color:#000
    classDef data fill:#fbbf24,stroke:#92400e,color:#000

    class R client
    class M,M2 service
    class D,APP async
    class EX service
    class ANS data

Three things to hold onto. The model never executes anything โ€” it emits a name and a JSON argument object, and your code decides whether to run it. It always takes at least two calls, because the model needs a second pass to see the tool result. And you append both the assistantโ€™s tool-call message and your tool result message to history, matched by call id, or the next call will be incoherent.

The hard parts โ€” schema design, argument validation, parallel calls, loop termination, and the authorisation question of what a model-chosen tool is allowed to touch โ€” are the subject of Tool Calling.


Multi-Turn Assembly

def build_messages(conversation_id: str, user_text: str, retrieved: list[str]):
    system = [{"role": "system", "content": SYSTEM_PROMPT}]          # from your code
    history = load_recent_turns(conversation_id, budget_tokens=4000)  # from your DB
    context = (
        [{"role": "user", "content": "Reference material:\n" + "\n---\n".join(retrieved)}]
        if retrieved else []
    )
    return system + history + context + [{"role": "user", "content": user_text}]

load_recent_turns takes a token budget, not a turn count, because ten long turns can exceed forty short ones. The ordering matters as well: system first, then history, then retrieved context, then the live question. Retrieved material sits closer to the question so it reads as evidence for this turn rather than as something the user said earlier. Every part of that list competes for the same context window, which is the budgeting problem in Tokens and Cost Math.


The Error Surface

This table is the part of the page most worth memorising. The failure modes are routine, not exceptional, and the correct response differs sharply per class.

Error class What it means Correct handling
Rate limited You exceeded requests or tokens per interval Retry with exponential backoff and jitter, honouring any retry_after. Longer term, add client-side admission control and queue non-interactive work โ€” see Rate Limiting for what the provider is enforcing
Quota or billing exhausted Hard account limit, not a transient throttle Do not retry. Alert an operator. Fail the request with a clear message
Context length exceeded Assembled input plus requested output exceeds the window Not retryable unmodified. Trim history, summarize, or retrieve less. Better: budget before sending so this never fires
Client timeout No response inside your deadline Retryable, but assume the call may have completed and billed. Set deadlines above observed P99 and prefer streaming to distinguish a stall from slow work
Transient 5xx or overloaded Provider-side capacity or fault Retry with backoff, then fail over to another model or provider, then open a circuit breaker
Invalid request Bad parameter, malformed schema, unsupported field Never retry unchanged โ€” it is a bug in your code and will fail identically forever
Authentication failure Key revoked, expired, or wrong project Never retry. Page someone. See Authentication
Content filter or refusal Input or output tripped a safety policy Not retryable as sent. Return a product-level message, log the case for your eval set, and never surface the raw provider error
Valid response, unusable output HTTP 200 with malformed JSON or a wrong-shaped answer One repair turn feeding the validation error back, then a deterministic fallback. Count these as a first-class error rate

That last row is the one teams forget to instrument. A 200 response that fails schema validation is a failure your users feel and your HTTP dashboards will not show. Track it next to your 5xx rate.


Idempotency and Retry Safety

Retrying a normal GET is free. Retrying a model call is not, and for two separate reasons: you may be billed twice, and because generation is non-deterministic you may get a different answer the second time. A timeout is the dangerous case โ€” the request may have succeeded on the provider side while you never saw the response.

The fix is to make retries safe at your own boundary rather than hoping the provider does it for you:

Background on both halves: Idempotency and Retry and Backoff.


Provider Abstraction: Wrapper or Framework

Bad โ€” raw SDK calls scattered everywhere. Forty call sites, each with its own timeout, none with retries, model name hardcoded in twelve of them. Adding token logging means touching forty files, and a provider outage has no single place to fail over from.

Good โ€” one thin internal client. A single module every call site goes through, owning timeouts, retries, model selection from config, token accounting, structured logging, and error normalisation into your own exception types. This is a few hundred lines and pays for itself the first time you need a fleet-wide change. It is ordinary API design applied to an outbound dependency.

Great โ€” that wrapper plus a deliberately narrow interface. Keep the surface to roughly โ€œgiven messages and options, return text or a tool call, streaming or not.โ€ That is small enough to implement against a second provider in an afternoon, which is what makes failover and side-by-side model selection real options rather than aspirations. Resist widening it to cover every provider-specific flag; expose an escape hatch for the one caller that needs one.

A framework earns its place when you want prebuilt connectors and loaders for a prototype, or when you genuinely need agent orchestration and durable multi-step state. It is overkill when your production path is one well-understood call and you are importing a large dependency to avoid writing thirty lines โ€” you inherit its abstractions, its upgrade cadence, and an extra layer between you and the error messages you need to read.


When to Use

โœ… Call the API directly, through your own thin client, when:

โŒ Reach for something heavier when:


Common Interview Questions

Q1: Why does a chat API make you resend the entire conversation on every request?

Because the endpoint is stateless โ€” it holds no session and no server-side thread. The model is a function from the message list to the next message, so the message list is the only state there is, and it lives in your database. The consequence is that cost and latency grow with conversation length, since turn twenty pays to re-read turns one through nineteen. That is why history trimming, summarization, and retrieval instead of stuffing are not optimisations but requirements for long-lived chats.

Q2: A request times out. Do you retry it, and what could go wrong if you do?

You can retry, but only knowing that the call may have completed and billed on the provider side while the response never reached you. A naive retry risks double spend and, because generation is non-deterministic, a different answer that can produce two divergent records of the same turn. Guard it with a request key derived from the inputs, persisted before the call: on retry, check the key and return the stored result if one landed. And if the call triggers side effects through tools, put the idempotency guarantee in the tool rather than in the completion.

Q3: When would you stream and when would you not?

Stream when a human reads the output as it appears, because time to first token dominates perceived latency and streaming also lets a user cancel and stop paying for further tokens. Do not stream when a program consumes the output โ€” a partial JSON object is unusable, so you would buffer it yourself anyway, and you lose the ability to validate before acting. The operational cost of streaming is real: long-lived connections, disconnect handling, and usage numbers that only arrive in the final chunk.

Q4: How do you handle rate limits properly beyond adding a retry?

A retry with backoff and jitter is the reactive floor. The real fix is not exceeding the limit: bound your concurrency to what your quota actually allows, since concurrency above it converts into 429s rather than throughput, and push non-interactive work such as evals and bulk jobs onto a queue that drains at a controlled rate. Separate interactive traffic from batch traffic so a backfill cannot starve live users, and distinguish a transient throttle from hard quota exhaustion โ€” the first is retryable, the second should page someone.

Q5: Is a wrapper around a provider SDK worth writing, or is it premature abstraction?

A thin one is worth it almost immediately, because it gives you a single place for timeouts, retries, model configuration, token accounting, and error normalisation. Without it, a fleet-wide change means touching every call site and a provider outage has nowhere to fail over from. The discipline is keeping the interface narrow โ€” messages in, text or tool call out, streaming optional โ€” because a narrow interface is what makes a second provider cheap to add. It becomes premature the moment you start reimplementing a frameworkโ€™s chain and agent abstractions you do not need.


Build it in code: Agentic AI Course · Fundamentals: Core Concepts

Free system design + DSA prep. If it helped you crack an interview, consider supporting.

SensAI SensAI
Beta
Listening...
Tap mic to stop voice mode

Shape what we build next

Every piece of feedback is read by the team and directly influences our roadmap.

What type of feedback?

Install SystemCraft

Add to your home screen for instant access, offline reading, and a distraction-free experience.

Offline reading Faster loads No browser tabs App-like feel

Unlock AI Features

One click to activate - no payment, no credit card. Just sign in and you're in.

AI code review and hints
SensAI chat assistant
AI mock interviews
Whiteboard analysis
100% free during early access