Recommended Free Tools
Short answer: Cohere released Rerank 4 on December 11, 2025, with rerank-v4.0-pro and rerank-v4.0-fast. Both list a 32,768-token context, versus 4,096 tokens for Rerank 3.5—an eightfold nominal increase, not fourfold. The larger window can reduce chunk-boundary errors in long enterprise records and improve the context given to RAG systems and agents, but it does not guarantee better rankings or fewer agent failures. Test it on your own corpus before migrating.
What Cohere launched
Rerank 4 is a second-stage relevance model: your lexical, vector, or hybrid search first retrieves candidates, then Rerank scores each candidate against the query and reorders the list. Only the best results are normally sent to a generator or agent. It does not replace the search index, embedding model, access-control layer, or language model. See Cohere’s reranker overview.
The release has two variants. rerank-v4.0-pro targets maximum ranking quality and complex queries; rerank-v4.0-fast targets lower latency and higher throughput. Cohere lists both as multilingual and able to process text and serialized semi-structured data such as JSON. Support for more than 100 languages does not imply equal quality for every language, domain, or mixed-language query.
Rerank 4 versus Rerank 3.5
| Capability | Rerank 3.5 | Rerank 4 |
|---|---|---|
| Context length per query-document evaluation | 4,096 tokens | 32,768 tokens |
| Variants | One listed model | rerank-v4.0-pro and rerank-v4.0-fast |
| Languages | Multilingual | Multilingual |
| Structured data | JSON and semi-structured data | JSON and semi-structured data |
| Document handling | Automatic chunking at roughly 4K-token context | Automatic chunking at roughly 32K-token context |
| Positioning | General enterprise reranking | Pro: highest quality; Fast: lower latency and higher throughput |
These figures come from Cohere’s model table, reranking guidance, and release notes.
#1 Best Overall
Is the increase fourfold or eightfold?
The “quadruples” wording comes from launch coverage, including VentureBeat’s headline. Cohere’s official limits are 32,768 versus 4,096 tokens, which is eight times the nominal context length. Usable document space is lower because the query and reserved tokens share that budget, so a claim about a fourfold usable-document increase may reflect different assumptions. For engineering decisions, use the documented limits and measure your requests.
What the 32K context changes
More complete inspection of long documents
Cohere says Rerank 4 can use chunks up to 32,764 tokens after reserved tokens, compared with about 4,093 for 3.5. A policy, contract, manual, ticket history, email thread, or code file that previously required several chunks may fit into one evaluation. Cohere estimates 32,768 tokens at roughly 48–50 pages, but page counts vary with tables, code, formatting, and tokenization.
Fewer chunk-boundary mistakes
With a small window, a definition can be separated from its exception, a table row from its heading, or a contract clause from the effective date that qualifies it. A larger window reduces this loss of context. It does not remove it: documents above the effective limit are still chunked, and poorly chosen ranking units can still hide the relevant passage.
Better handling of enterprise records
Likely beneficiaries include legal and procurement documents, support and incident histories, internal policies, product manuals, technical documentation, email threads, JSON or YAML-like records, and source code. Rerank 4 can compare a query with more of each candidate, which is especially useful when the relevant evidence is not near the beginning.
Free tools Windows power users keep installed
One-click scans. No signup required.
Potentially cleaner agent context
If reranking places fewer, more relevant records at the top, an agent may spend fewer tokens on irrelevant evidence and have fewer opportunities to follow an unrelated instruction. That is a plausible retrieval-stack effect and Cohere’s product positioning—not a published, universal percentage reduction in agent errors. Planning mistakes, prompt injection, bad tool schemas, permission errors, state bugs, outages, and generator failures remain separate risks.
How query and document tokens are counted
The context limit is for the query plus one document evaluation; it is not the total number of tokens you can submit across every document in a request. In Rerank 4, a query can use up to 16,384 tokens. Longer queries are truncated to the first 16,384 tokens. In Rerank 3.5, the corresponding limit is 2,048 tokens. Remove irrelevant conversation history before reranking so the actual information need gets the budget.
Documents exceeding the combined limit are automatically chunked. Cohere’s examples describe taking the maximum relevance score across chunks, which creates a “best passage wins” result rather than a holistic judgment of the entire record.
Request-size ceiling
Cohere documents an error when the effective document-and-chunk count exceeds 10,000. The practical constraint is:
number of documents × max_chunks_per_doc ≤ 10,000
The default max_chunks_per_doc is 1. Raising it makes long-document coverage possible but reaches the ceiling sooner; a 32K context does not permit unlimited long documents in one call.
When to keep application-level chunking
Let Rerank 4 handle larger passages when preserving context is the goal, but chunk deliberately when:
- a record contains unrelated subjects;
- citations must point to an exact section;
- metadata or access filters apply per section;
- one file contains conflicting versions or effective dates;
- the document is substantially larger than 32K tokens; or
- a whole-record score would obscure the relevant subsection.
For JSON or YAML, serialize records consistently and put high-value fields early when truncation is possible. Cohere notes that key order affects what survives truncation.
Choosing Pro, Fast, or 3.5
| Choose | Good fit | Qualification |
|---|---|---|
rerank-v4.0-pro |
Nuanced queries, high-value results, legal, financial, medical, policy, or risky agent actions | Positioned for quality; measure latency and cost on your workload |
rerank-v4.0-fast |
Interactive search, large candidate sets, high-throughput services, repeated agent loops | Positioned for speed and throughput; measure any quality trade-off |
rerank-v3.5 |
Short, clean documents; proven pipelines; strict latency budgets | Still a sensible baseline when 4K context is sufficient |
Cohere does not publish a universal latency multiplier or accuracy delta between Pro and Fast. Route different workloads to different variants rather than assuming one model fits every path.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Repair Windows errors before they cause bigger problems3Fix the driver behind crashes, sound loss and screen glitchesRank #4
Minimal API migration
The model identifier is the main change in an existing Cohere integration. Cohere’s V2 Python example is:
import cohere
co = cohere.ClientV2()
query = "What is the company's parental leave policy?"
documents = [
"Document text or retrieved passage 1",
"Document text or retrieved passage 2",
"Document text or retrieved passage 3",
]
response = co.rerank(
model="rerank-v4.0-pro",
query=query,
documents=documents,
top_n=5,
)
for result in response.results:
print(result.index, result.relevance_score)
Keep candidate generation, filters, document IDs, and authorization unchanged while comparing models. Scores are for ordering and locally calibrated thresholds; they are not universal probabilities of correctness.
A defensible migration test
- Preserve the current
rerank-v3.5path and assemble representative queries, relevant document IDs, labels, final answers, and agent outcomes. - Run identical candidates through 3.5,
rerank-v4.0-fast, andrerank-v4.0-pro. - Measure candidate recall, NDCG or MRR, final top-k precision, answer faithfulness, citation correctness, agent task completion, P50/P95 latency, and cost per query.
- Manually inspect disagreements involving long documents, exceptions, tables, JSON, multilingual queries, and similar records with different effective dates.
- Choose a variant per route, then release behind a feature flag or percentage split.
- Keep rollback to 3.5 until production evaluation is complete.
A reranker cannot recover a document absent from the first-stage results. Also check stale indexes, metadata filters, OCR and parsing quality, query rewriting, and access-control boundaries before attributing a failure to model quality.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Cost, deployment, and availability
Cohere offers a hosted API, Model Vault, private VPC or on-premises deployment, and cloud integrations including Azure AI Foundry and Oracle OCI. See Rerank, Model Vault, Oracle’s OCI documentation, and Azure AI Foundry.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallBest Value
Cohere’s pricing page says API usage is charged by searches and dedicated Model Vault deployments by instance. The page listed these Model Vault signals in August 2026:
| Deployment | Listed rate |
|---|---|
| Rerank 4 Fast, Medium | $5/hour; $3,250/month |
| Rerank 4 Pro, Medium | $5/hour; $3,250/month |
| Rerank 4 Pro, Large | $10/hour; $6,500/month |
These are published signals, not a quote: enterprise contracts, marketplace terms, minimums, and regional taxes can differ. See Cohere pricing. Larger inputs may require more inference work, so fewer chunks do not automatically mean lower cost.
Alternatives to evaluate
| Option | Why consider it | Trade-off |
|---|---|---|
| Voyage AI | Another managed reranking provider or an existing Voyage stack | Requires separate quality, latency, compliance, and deployment testing |
| Jina AI | Multilingual and developer-oriented retrieval | Fit depends on your corpus and operating requirements |
| Mixedbread or BAAI BGE | Open-weight or self-hosted deployment and data control | You operate serving, scaling, monitoring, upgrades, and security |
| Elasticsearch-native ranking | Tight lexical, filtering, and operational integration | Not the same cross-attention behavior as a dedicated neural reranker |
| ZeroEntropy | Another commercial API to benchmark | Procurement and deployment footprint must be assessed |
Current prices and availability for these alternatives were not established here; obtain current vendor terms before buying.
Verdict
Rerank 4 is a meaningful upgrade when relevant evidence is routinely buried beyond a few thousand tokens, split across tables or exceptions, or embedded in long multilingual and structured records. Start with Fast for measured throughput-sensitive routes and Pro where ranking mistakes are expensive. Stay on 3.5 when documents are short, quality is already sufficient, or the real defect is candidate recall, filtering, parsing, or authorization. The context headline is a reason to test—not proof of better enterprise accuracy or safer agents.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




