Docker Model Runner (DMR) documents OpenAI-compatible, Anthropic-compatible, and Ollama-compatible API formats. To test whether your application works with each one, hold the model and prompt intent constant, send a request to each API’s own endpoint, and assert the response contract your application relies on. Do not expect identical schemas, behavior, or generated text: Docker documents differences between the formats.
What this test can—and cannot—prove
A contract test checks whether an API response satisfies the requirements of your application. For a chat feature, those requirements might include a successful HTTP response, a parseable response in the expected format, non-empty assistant text, and any fields your code needs to consume.
As an Amazon Associate I earn from qualifying purchases.
Running the same model through three interfaces can reveal whether your application’s assumptions hold for each format. It does not establish that the APIs are interchangeable or that they produce identical outputs. Keep assertions focused on response structure and application behavior rather than exact generated prose.
Set a controlled baseline
Before sending requests, define the behavior under test—for example, one non-streaming chat turn with a user message and a maximum token bound. Record the expected status, format-specific response fields, and application-level output constraints. If your application uses streaming, tools, or structured output, make those separate test cases rather than assuming a basic chat test covers them.
#1 Best Overall
Use the same DMR model identifier in each request. Docker’s API reference gives examples including ai/smollm2 and the tagged identifier ai/smollm2:360M-Q4_K_M. For reproducibility, capture the model ID or tag, Docker Desktop or Docker Engine version, host operating system, inference engine, context configuration, sampling settings, and hardware backend alongside test results.
Configure each API client for its documented route
The base URL and endpoint depend on the client family. Docker’s examples use these addresses and routes; use the documented host TCP setup and enable TCP access where applicable. These are local addresses intended for a client running on the same host.
Rank #2
| API format | Base URL | Chat route | Contract to assert |
|---|---|---|---|
| OpenAI-compatible | http://localhost:12434/engines/v1 |
/chat/completions (full path: /engines/v1/chat/completions) |
OpenAI-style chat-completions response fields used by your application |
| Anthropic-compatible | http://localhost:12434 |
/v1/messages |
Messages response fields used by your application |
| Ollama-compatible | http://localhost:12434 |
/api/chat |
Ollama chat response fields used by your application |
For Ollama-compatible prompt completion rather than chat, Docker documents /api/generate. Do not normalize the requests into a made-up shared schema: retain each API’s actual payload shape and validate its own response envelope. See Docker’s Model Runner API reference for the documented routes and request formats.
Recommended Free Tools
Build the three contract checks
OpenAI-compatible chat completions
Send the request to http://localhost:12434/engines/v1/chat/completions using the payload shape expected by the OpenAI-compatible endpoint. Include the same model identifier and equivalent user-message intent as in the other calls. Test only parameters your application actually sends; Docker lists parameters including model, messages, max_tokens, temperature, top_p, streaming, stop, and penalty parameters.
Rank #3
Anthropic-compatible Messages
Send a Messages request to http://localhost:12434/v1/messages. Keep the model and prompt intent aligned with the baseline, but preserve the Messages schema. If the application depends on a system prompt, streaming, or stop sequences, test those behaviors explicitly using the corresponding format’s fields.
Ollama-compatible chat
Send a chat request to http://localhost:12434/api/chat and validate the Ollama-format response fields your application consumes. Use /api/generate instead only if the application uses prompt completion rather than chat.
Apply format-specific assertions
For each call, check HTTP status, successful parsing, the expected API-specific response envelope, and your application’s output constraints. A practical test suite can share the model identifier and behavioral intent while keeping separate request builders and validators for each API. Label this as a test design, not as a claim that Docker supplies a first-party contract-testing suite.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Test differences that can break application assumptions
Authentication
Do not treat an OpenAI-style Authorization header as protection for DMR: Docker documents that the OpenAI-compatible implementation ignores it. Docker also states that the Model Runner API is not authenticated. Do not expose the API to untrusted networks during testing. See the Docker Model Runner overview for its networking and authentication notes.
Best Value
- Docker, Docker Swarm, Docker Compose, Programmer, Developer, Coding, Programming, Software Engineer, Code, DevOps, Deploy, Deployment, Kubernetes, Salt, Puppet, Chef, Terraform, Container, AWS, Azure, Cloud, Geek, Funny, Computer, Software, Tech, IT
- Integration, Scrum, Compile, Compilation, Science, Bug, Debug, Python, Linux, Java, Javascript, Scala, Dotnet, Kotlin
- Lightweight, Classic fit, Double-needle sleeve and bottom hem
Function calling
If your application uses function calling, make it a separate test with a compatible model and the documented llama.cpp conditions. Docker describes this support as conditional; a passing basic chat request does not establish that tool or function calls will work.
Token counts
Do not assert that DMR token counts match OpenAI’s. Docker says token counting uses the model’s native encoder, which may differ from OpenAI’s.
Streaming and errors
If your application streams responses, check event framing, parsing, and completion behavior for each API independently. Also exercise the error cases your application must handle and validate the format-specific error response rather than assuming one schema applies to all three. Docker provides streaming examples, but those examples do not establish identical streaming semantics across formats.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Make the test environment reproducible
The model and API format are not the only variables: the inference engine and host platform can affect what is available and how a run behaves. Docker documents llama.cpp as the default engine; vLLM and Diffusers have narrower platform and GPU support. Record engine and platform details with each test run, and consult the Model Runner requirements and setup documentation for current availability and prerequisites. Docker’s documented minimum Docker Desktop versions are Windows 4.41+ and macOS 4.40+; verify the requirements for your environment because these specifications can change.
Docker says Testcontainers for Java and Go and Docker Compose support Model Runner, which can help create repeatable test environments. The existence of that support is not itself a ready-made suite for comparing API contracts; your tests still need to define and assert the behavior your application requires.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




