Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Scan for outdated or missing drivers - takes under a minute3Clear out junk files and repair common Windows errorsSet Gemini API generation options for the specific model you call: treat maxOutputTokens as a hard ceiling, keep Gemini 3’s temperature at its recommended default of 1.0, and choose safety thresholds with your application’s risks in mind. For reasoning-capable models, the output cap also covers thought tokens, so a low limit can leave you with a partial or empty response.
How to set Gemini API generation options
Generation options go in the request’s GenerationConfig. Google’s API reference lists fields such as maxOutputTokens, temperature, topP, topK, candidate count, stop sequences, and response MIME type. The available options and their defaults can depend on the model. Check the selected model’s documentation before relying on a setting or copying a generic configuration: GenerateContent API reference.
Keep field names consistent with the SDK you use. The API reference uses camel case such as maxOutputTokens; some guide prose and examples use snake case, such as max_output_tokens. Use the naming convention required by your SDK rather than mixing them.
Choose a maximum output-token limit
maxOutputTokens limits the number of tokens included in a response candidate. It is a ceiling, not a requested answer length. The default and maximum vary by model, so look up the selected model’s output_token_limit instead of assuming one universal value. Leave enough room for the answer you need, while staying within that model’s supported limit.
#1 Best Overall
Account for thought tokens
For thinking-capable models, the output-token cap includes thought tokens as well as the answer. A small cap can stop generation while the model is reasoning; the result may be truncated or empty, with MAX_TOKENS as the finish reason. If you want to reduce cost or latency without imposing a very low hard ceiling, Google’s thinking guide recommends adjusting thinking_level rather than sharply limiting max_output_tokens: Gemini thinking guide.
Choose temperature for the model and task
Temperature affects the randomness of sampled output. Its default and supported range are model-dependent. Google’s API reference describes a general range of 0.0–2.0, while its troubleshooting guidance lists 0.0–1.0 among parameter checks. These differing reference contexts are not a universal range for every model and endpoint. Validate the value against the model and API path you actually use.
Rank #2
For Gemini 3, keep the default at 1.0
Google strongly recommends leaving temperature at its default of 1.0 for all Gemini 3 models. The Gemini 3 guide warns that changing it, particularly setting it below 1.0, can cause unexpected behavior such as looping or weaker performance on complex math and reasoning tasks: Gemini 3 developer guide. Do not assume that lowering temperature will make Gemini 3 reliably deterministic. If you experiment with sampling settings for another model or task, test the actual outputs you plan to use.
Set safety thresholds per request
Gemini’s safety settings let you choose block thresholds for four harm categories. The threshold indicates the probability level at which content is blocked:
Rank #3
| Threshold | What it blocks |
|---|---|
BLOCK_ONLY_HIGH |
High-probability content |
BLOCK_MEDIUM_AND_ABOVE |
Medium- and high-probability content |
BLOCK_LOW_AND_ABOVE |
Low-, medium-, and high-probability content |
OFF and BLOCK_NONE |
Additional threshold options listed in Google’s guide; check the current documentation for their behavior and availability for your model. |
The adjustable categories in Google’s guide are harassment, hate speech, sexually explicit content, and dangerous content. The guide describes harassment as negative or harmful comments targeting identity or protected attributes; hate speech as content it characterizes as rude, disrespectful, or profane; and dangerous content as material that promotes, facilitates, or encourages harmful acts. Consult the safety settings guide for current category definitions and configuration details.
If you do not set a threshold, Google says the default block threshold is Off for Gemini 2.5 and Gemini 3. Do not apply that default to other model families without checking their current documentation. Threshold availability and parameter support can change by model.
Rank #4
Handle safety blocks in application code
Safety settings are passed with a request. Google says content receives a category and probability rating. Inspect the response rather than assuming every request produces user-facing text:
- For a blocked prompt, check
promptFeedback.blockReason. - For a response candidate, inspect
finishReasonandsafetyRatings. - A safety-blocked candidate uses the
SAFETYfinish reason, and the blocked content is not returned.
Use those fields to distinguish a safety block from an ordinary answer and provide an appropriate application response, such as asking for a permissible reformulation or explaining that the request could not be completed. Test the handling with realistic safe and unsafe examples for your own use case. Making thresholds more permissive may reduce blocks on borderline content, but increases the need for application-level review; simply disabling filters is not a substitute for testing.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Best Value
Validate settings when a request fails
If the API rejects a generation option, check that the model and API version support it and that the value is valid for that endpoint. Google’s troubleshooting guide specifically advises checking version and model feature support when parameters cause errors: Gemini API troubleshooting. Model limits, parameter support, and safety defaults are not interchangeable across the Gemini API.
Safety settings are only one safeguard
Filters are not a guarantee that output will be accurate, unbiased, or harmless. Google advises developers to assess risks for their application, consider mitigations, conduct appropriate safety testing, gather feedback, and monitor use: Safety guidance. Treat thresholds as one control in that broader process, not as fact-checking or a replacement for application-specific safeguards.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




