Test an AI API integration at three distinct boundaries: verify the requests and responses your code depends on, exercise application workflows with deterministic test doubles, and evaluate whether real model outputs still meet product requirements. A successful HTTP response proves none of the latter two. Keep provider and model behavior tests separate from API-contract tests so a schema change does not get confused with a probabilistic change in model behavior.
What counts as a breaking change?
A break can occur in the API contract, the SDK’s conversion of your application data into provider requests, the transport or streaming behavior, or the model behavior your product relies on. These are related but not interchangeable. A response can remain valid under the API schema while becoming less useful for a particular task; conversely, model behavior can remain acceptable while an SDK upgrade changes a public interface.
OpenAI’s API Reference describes some additive changes—such as new optional request parameters and response properties, or a different property order—as backward compatible. That does not mean your application should ignore all changes. Test the specific required fields, types, values, and behaviors your code depends on, rather than rejecting every unfamiliar field or assuming object-key order is meaningful. OpenAI also cautions that prompting behavior can change between model snapshots, so schema compatibility and behavioral consistency need separate tests.
This guidance uses OpenAI documentation as a concrete example, not as a guarantee for every provider. For a multi-provider product, apply the same test layers to each provider’s own API, SDK, release, and deprecation policies.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Build a test suite around the boundaries
Choose each test by asking what it can actually prove. A test double can establish that your workflow handles a scripted tool call; it cannot prove that the provider accepts the request payload your SDK emits. An evaluation can reveal a quality regression; it does not establish that authentication headers or streaming events are correct.
| Test layer | What it can establish | What it cannot establish by itself |
|---|---|---|
| Contract and serialization | Your application sends required fields and handles the response and errors it relies on; with a real adapter and controlled transport, it can also inspect serialized requests and transport details. | That a model response meets product quality requirements. |
| Deterministic workflow tests | Routing, retries, state transitions, tool handling, output processing, and failure branches under scripted conditions. | Provider request conversion, authentication, actual HTTP or WebSocket payloads, or provider-specific streaming chunks when those are not exercised by the test. |
| Live integration checks | Selected provider-side paths that a controlled transport cannot faithfully exercise, such as actual authentication or provider-managed lifecycle behavior. | Broad behavioral quality across prompts, data, and model changes. |
| Model evaluations | Whether representative outputs meet task-specific criteria for a model and configuration. | Correct request serialization or transport behavior. |
OpenAI’s Agents JavaScript SDK testing guide describes in-memory doubles for scripted responses and workflows, while explicitly excluding provider request conversion, HTTP or WebSocket payload details, authentication headers, provider-specific stream chunks, and full provider lifecycle fidelity. Keep a test double at the boundary it models; add real-adapter coverage for the boundaries it omits.
1. Specify the contract your application relies on
Write down the subset of each request and response that matters to your integration. Assert required fields, expected types, permitted values, tool or function schemas, and error behavior. Avoid brittle assertions on incidental details such as property order or opaque identifiers unless your own code genuinely depends on them.
- For requests, check the endpoint and required fields, the selected model or model identifier, configuration values, and tool definitions your application expects to send.
- For responses, verify the fields and types your code consumes, including tool-call arguments and the handling of missing, malformed, or partial content.
- For errors, cover the cases that change application behavior: whether to retry, fall back, stop, or surface a failure.
- For tool use, test both valid arguments and schema-validation failures. Successful JSON parsing is not proof that arguments satisfy your application’s contract.
Do not assume a schema is enforceable merely because strict mode is enabled. OpenAI documents that strict schema enforcement applies only to supported model and configuration combinations and supported JSON Schema subsets. Validate your supported schema against those constraints and test the behavior when a schema is rejected.
Free tools Windows power users keep installed
One-click scans. No signup required.
Rank #2
2. Test application workflows deterministically
Use fixed model responses or scripted tool calls to test the application’s control flow without requiring a live model request for every application-level case. This makes it practical to test the branches that are easy to miss in a happy-path integration test.
- Route a response to the right handler and verify the resulting state transition.
- Script a multi-turn tool loop, including argument handling and the result returned to the next turn.
- Exercise retries, timeouts or failures as represented at your chosen abstraction boundary, and verify that retry limits and fallbacks behave as intended.
- Feed malformed, incomplete, or unexpected content into output handling and check that the application fails safely rather than treating it as a valid result.
- Use a changed scripted sequence to detect workflow drift when an application path no longer reaches the expected state.
The OpenAI Agents JavaScript SDK documents recipes for fixed responses, multi-turn tool loops, streaming, model failures, and workflow-drift detection. Those tests are useful precisely because they are controlled; they should not be reported as proof of provider wire compatibility.
3. Exercise the real adapter and transport
To catch incompatibilities between your code and a provider integration, run the real provider adapter against a controlled or mocked network transport. Inspect the request the adapter produces and test the transport-level behavior your application depends on. This keeps the test repeatable while covering more than a workflow double can.
- Check serialization, endpoint selection, and required headers, including authentication handling without exposing secrets in test output.
- Exercise relevant HTTP success and error paths and verify how the adapter maps them into application-visible results.
- For streaming, validate provider-specific events and chunk handling rather than assuming a generic sequence of text fragments.
- Keep a smaller set of live integration tests for paths that a controlled transport cannot faithfully represent, such as confirming credentials work or exercising a provider-side lifecycle.
Scope live tests deliberately: they can catch real-environment failures, but they are not a substitute for fast deterministic workflow tests or model evaluations. Separate test credentials and environments from production, and make the provider, endpoint, and model explicit in test configuration so a run’s result can be interpreted correctly.
Rank #3
- Contains one (1) API 5-IN-1 TEST STRIPS Freshwater and Saltwater Aquarium Test Strips 25-Count Box
- Monitors levels of pH, nitrite, nitrate carbonate and general water hardness in freshwater and saltwater aquariums
- Dip test strips into aquarium water and check colors for fast and accurate results
- Helps prevent invisible water problems that can be harmful to fish and cause fish loss
- Use for weekly monitoring and when water or fish problems appear
4. Evaluate model behavior separately
Maintain representative evaluation cases for the product’s actual tasks. Score requirements that matter to users—for example, answer correctness, required output structure, tool selection, refusal or guardrail behavior—rather than treating a completed request as success. OpenAI’s evaluation guidance characterizes evals as structured measurements and recommends them because generative outputs vary.
Run the suite against the current and proposed model or configuration when changing either. Review score changes alongside representative output differences; an aggregate score can identify a regression, while examples help show what changed and whether it matters to the application.
Do not substitute an industry benchmark or a generic numerical measure for an application evaluation. OpenAI distinguishes industry benchmarks, scoring measures, and tests built for a particular application. Select criteria and examples that correspond to your own user-facing task, and keep the evaluation dataset and scoring method stable enough to compare runs.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Make failures reproducible and upgrades reviewable
Attach enough context to every failure report that another engineer can reproduce and classify it. Record the provider and endpoint, SDK and version, model identifier or pinned snapshot, relevant configuration, test or evaluation dataset version, and the failing request or case with secrets removed. Label whether the failure came from a contract check, deterministic workflow, transport integration, or model evaluation.
Pin model snapshots when repeatable prompting behavior matters, and use evaluations when changing snapshots. Pinning does not remove the need to test; it makes the tested model target explicit. Treat SDK versioning independently: OpenAI’s Python Agents SDK documents a modified 0.Y.Z policy under which minor releases can include breaking public-interface changes, and its release guidance recommends pinning 0.0.x when avoiding breaking changes. Do not infer SDK guarantees from the provider API’s compatibility policy.
For each upgrade, review the SDK’s own release notes and breaking-change guidance, run the relevant test layers, and compare evaluation results where model behavior is implicated. Use provider changelogs and deprecation notices to plan API or model migrations before a removal reaches production.
Plan for model and service deprecations
OpenAI’s current Deprecations documentation says generally available models normally receive at least six months’ notice before retirement, while specialized generally available variants normally receive at least three months. Preview models can receive much shorter notice, and exceptions may apply for safety or compliance. These are OpenAI’s stated timelines, not a universal industry policy; check the relevant provider notice for each service you use.
As of October 4, 2026, OpenAI’s documentation schedules its Evals content to become read-only on October 31, 2026, and the dashboard and API to shut down on November 30, 2026. The same documentation points to Promptfoo as a migration path. If your team relies on that platform, confirm the current migration details and preserve the datasets and results you need before those dates. This timeline is time-sensitive and should be checked against OpenAI’s current notice when acting on it.
Recommended Free Tools
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




