To make a retrieval-augmented generation (RAG) pipeline easier to change, keep its use cases dependent on application-owned ports—not provider SDKs—and put each model, retrieval system, or storage integration behind an adapter. This lets you replace infrastructure without rewriting the core workflow, provided you explicitly account for differences in provider capabilities and data formats.
What hexagonal architecture means for a RAG pipeline
Hexagonal architecture, also called ports and adapters, separates an application’s core behavior from the external systems it uses. AWS Prescriptive Guidance describes ports as technology-agnostic entry points into an application component. A port defines an interaction the application needs; an adapter implements that interaction for a particular technology and translates between its formats and the application’s types.
As an Amazon Associate I earn from qualifying purchases.
For RAG, the core might accept documents, prepare chunks, request embeddings, index content, retrieve evidence, assemble context, and generate an answer. It should express those steps in terms of application concepts rather than calls to a model or vector-database SDK. A hosted embedding API and a local embedding model can then be alternative adapters for an embedding port, for example.
How the parts fit together
Requests can enter through a driving adapter, such as an HTTP API or a queue consumer. The adapter calls an application use case, which coordinates work through ports. Outbound adapters connect those ports to models, indexes, and other services. A composition root—the place where the application is configured—selects the concrete adapters for a deployment or test.
#1 Best Overall
Driving adapters: HTTP API | CLI | queue | scheduled ingestion
|
v
Application use cases: IngestDocuments | AnswerQuestion | ReindexCollection
|
application-owned ports
/ |
v v v
Embedder IndexWriter AnswerGenerator
| | |
v v v
model adapter retrieval model adapter
adapter
The diagram is conceptual: a particular system may combine or split these responsibilities. The important boundary is that application use cases depend on contracts the application owns, while provider-specific configuration, credentials, SDK objects, and response translation remain in adapters.
Which ports a RAG system may need
Start with the interactions that cross a real boundary or are likely to change. An interface for every helper function adds ceremony without necessarily protecting the application from meaningful change.
Rank #2
- Document source or ingestion input: supplies documents to the ingestion use case.
- Document transformer: parses, normalizes, or chunks content when those steps need an independently replaceable implementation.
- Embedder: converts text to vectors. Make model identity and vector dimension explicit; define whether the port supports single-item calls, batches, or both.
- Index writer: adds, updates, and deletes indexed content.
- Retriever: returns application-level documents and metadata, with any required filters or retrieval options represented deliberately.
- Answer generator: accepts a typed prompt or message request and returns a typed generation result.
- Optional reranker, clock, or telemetry ports: add these when the dependency is meaningful to the application and needs isolation or substitution.
Keep the boundary types small and stable. A document might contain text, a stable identifier, and metadata; a retrieval result might contain that document, a score, and source information. The adapter should map provider responses into those types rather than allowing provider response objects to leak into use cases.
Free tools Windows power users keep installed
One-click scans. No signup required.
Design contracts around capabilities, not a false promise of sameness
A shared interface reduces direct coupling; it does not make providers behave identically. Before defining a contract, write down what the application actually relies on. Then either require that behavior from every adapter or expose its availability as a capability or deployment choice.
Rank #3
| Boundary | Behavior to specify | What can vary by implementation |
|---|---|---|
| Embeddings | Model identity, vector dimension, batching expectations, and compatibility between document and query embeddings | Available models, request limits, and model-specific vector formats |
| Retrieval and indexing | Required filters, result shape, deletion behavior, pagination, and what a returned score means | Metadata-filter support, hybrid or sparse search, ranking semantics, and index features |
| Generation | Whether the use case needs streaming, structured output, tool calls, or a particular context capacity | Which features are supported and how refusals or safety signals are represented |
| Operations | Timeouts, cancellation, retryable error categories, and idempotency expectations | Rate limits, failure modes, authentication, retention, residency, and operational ownership |
Do not silently discard a filter an adapter cannot implement, treat scores from different retrieval systems as directly comparable, or promise identical streaming behavior when it is not supported. LangChain’s retrieval discussion describes different approaches, including similarity search, maximal marginal relevance, metadata filters, graph indexes, and retrievers built outside its vector-store approach. Those choices illustrate why the contract needs to match the application’s requirements rather than assume one universal retrieval behavior.
Test the use cases and each adapter at the right boundary
Ports make it possible to test application behavior without contacting a live provider in every unit test. Use fakes for fast tests of use-case decisions, then test each adapter against the contract the application depends on. A smaller integration suite can exercise real services.
Rank #4
- Use-case tests: verify workflow behavior with controlled fake documents, retrieval results, and generation responses.
- Adapter contract tests: check document and metadata mapping, expected dimensions, filtering, error translation, timeouts, and any streaming or structured-output behavior the application requires.
- Integration tests: confirm that the configured services work together in the deployed environment.
A passing contract test demonstrates that an adapter meets the tested expectations; it does not prove that two providers have identical answer quality or operational behavior. AWS Prescriptive Guidance identifies independent application testing and dependency mocking as benefits of the pattern, while also noting that adapters and additional layers bring maintenance overhead and can add latency.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchPlan provider changes as both code and data migrations
- Define the required behavior. Record the current port contract and the application capabilities in use, including filters, streaming, and error handling.
- Implement a new adapter. Translate the new provider’s requests and responses into application-owned types without changing the use case unless the requirement itself changes.
- Run contract tests and evaluations. Check adapter behavior, then compare retrieval and answer quality on representative queries and source documents.
- Plan index changes explicitly. If the embedding model or vector dimensions change, stored vectors may need re-embedding and the index may need re-indexing. An adapter cannot make incompatible vectors interchangeable.
- Choose a cutover and recovery approach. Base it on the system’s data and availability requirements; the architecture pattern alone does not guarantee a zero-downtime migration.
The reviewed sources do not quantify migration time or cost. The amount of work depends on the data, provider semantics, evaluation requirements, and operational plan.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Choose deployment options by operational fit
Provider-neutral ports can sit above different deployment categories, but deployment still determines who operates infrastructure, where data is placed, and how much control the team has. Google Cloud’s RAG architecture guide, last reviewed September 22, 2025, describes managed vector search, embeddings alongside operational data in AlloyDB, container-based RAG infrastructure, and a CI/CD architecture. These are categories to assess, not a comparative ranking.
| Deployment category | Questions to weigh |
|---|---|
| Managed vector search | Does the managed service provide the retrieval capabilities needed, and are its data placement and operating model acceptable? |
| Embeddings alongside operational data | Does keeping vectors with operational data fit the existing data architecture and retrieval requirements? |
| Container-based RAG infrastructure | Does the team need more control, and can it take responsibility for operating the components? |
| CI/CD-oriented architecture | How will changes to application code, adapters, indexes, and deployment configuration be validated and released? |
Compare actual candidates using the same workload and evaluation set. Consider supported capabilities, retrieval and generation quality, re-indexing effort, latency, reliability, privacy and data residency, operational work, portability, and cost. The cited architecture materials do not establish a universal winner or comparable benchmark results; verify current service details before choosing a deployment.
When the abstraction is worth its cost
Ports and adapters are useful when an application has multiple clients or integrations, a dependency may plausibly change, or provider isolation materially improves testing. A lighter design may be the better choice when there is one stable dependency, little domain behavior, and no meaningful replacement or testing requirement. The aim is to protect use cases from infrastructure decisions, not to build a second framework whose abstractions need their own maintenance.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →LangChain’s architecture documentation, generated and verified September 29, 2026, describes a three-layer structure of provider-agnostic core abstractions, orchestration, and partner packages that implement shared interfaces. It is an example of separating reusable contracts from integrations, not evidence that every RAG application should adopt that framework or that all its integrations support the same features. Its March 2023 retrieval article is useful as conceptual background; check current APIs before implementing against a framework.
Further reading
Hexagonal Architecture Explained: How the Ports & Adapters Architecture Simplifies Your Life, and How to Implement It, by Alistair Cockburn and Juan Manuel Garrido de Paz, is a general introduction to the pattern rather than a RAG implementation manual. Google Books lists the updated first edition as published by Humans and Technology Incorporated on April 15, 2025, at 196 pages.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




