What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Program-Aided Language Models (PAL) improve some kinds of LLM reasoning by splitting the work: the model interprets a question and writes a program, while an execution environment performs the calculation. The model then uses the result to answer in natural language. This can reduce arithmetic and procedural mistakes, but it cannot ensure that the model chose the right inputs, assumptions, or program.
What is a Program-Aided Language Model?
PAL is a method for combining a language model with a program runtime, commonly Python. It is not a separate model family. The LLM turns a natural-language task into executable steps; a runtime carries them out. The model remains responsible for understanding the question and designing the procedure, while the runtime handles deterministic operations.
The original PAL paper describes programs as intermediate reasoning steps and delegates their execution to a runtime such as Python. The preprint appeared on November 18, 2022; the work was published at ICML 2023. Read the paper or see its ICML publication page.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Clear out junk files and repair common Windows errors3Scan for outdated or missing drivers - takes under a minuteHow does PAL work?
- Interpret: The LLM identifies the quantities, rules, and desired output in the question.
- Generate: It writes a program that represents the calculation or procedure.
- Execute: A runtime runs the program and returns a result or an error.
- Respond: The LLM explains or formats the result for the user.
For example, consider: “A product costs $80, is discounted by 25%, and then taxed at 8%. What is the final price?” A PAL-style model might generate this illustrative Python snippet:
#1 Best Overall
- Use scikit-learn to track an example ML project end to end
- Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
- Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
- Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
- Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning
price = 80
discounted = price * (1 - 0.25)
final_price = discounted * 1.08
final_price
The runtime returns 64.8, which the model can present as $64.80. This example illustrates the pattern; it is not a reproduction of the original paper’s prompt.
Why can execution improve reasoning?
An LLM generates likely text; it is not inherently a calculator. It can describe a plausible sequence of arithmetic steps and still make a calculation error. PAL changes the division of labor: rather than asking the model to produce every intermediate result as prose, it asks the model to describe a procedure that an interpreter can execute.
- Deterministic computation: The runtime applies its language’s arithmetic and execution rules, reducing many token-prediction mistakes.
- State tracking: Named variables preserve intermediate values instead of relying on the model to keep them straight in prose.
- Reusable operations: Loops, functions, conditionals, and data structures can express multi-step procedures.
- Inspectability: Generated code can be logged, reviewed, tested, or rejected before it runs.
- Modular reasoning: The model handles language interpretation and planning; the runtime performs the specified computation. This is a form of neuro-symbolic cooperation, not fully symbolic reasoning.
In the original study, PAL was evaluated on 13 mathematical, symbolic, and algorithmic tasks. The authors reported that PAL using Codex scored 15 percentage points higher in few-shot GSM8K accuracy than PaLM-540B with chain-of-thought in their reported comparison. This is a historical result from a particular setup, not evidence that PAL universally outperforms larger models or improves every current model. See the paper’s benchmark context.
How does PAL differ from related approaches?
| Approach | Intermediate representation | Who performs the calculation? | Main strength |
|---|---|---|---|
| Chain-of-thought | Natural-language reasoning steps | The LLM | Flexible verbal decomposition |
| PAL | Executable program | An external runtime | Reproducible computation |
| Tool calling | Structured request to a specified tool | The selected tool | Access to APIs, databases, or actions |
| Retrieval-augmented generation (RAG) | Retrieved documents or passages | Usually the LLM, unless tools are also added | Grounding an answer in external information |
| Coding agent | Code, files, commands, and iterative actions | Multiple tools and runtimes | Broader software tasks |
PAL is not simply “chain-of-thought with Python”: its defining move is delegating computation to an execution environment. Tool calling is broader and need not express the full reasoning process as a program. RAG retrieves information; it does not, by itself, perform reliable calculations. A coding agent typically works across files and tools in a more open-ended loop than a focused PAL workflow.
What tasks are a good fit?
PAL is most useful when a task has a deterministic computational core, can be expressed clearly in code, and the inputs can be extracted reliably. Examples include:
Rank #2
- Arithmetic word problems, percentages, ratios, and unit conversion.
- Counting, combinatorics, symbolic algebra, and constraint checking.
- Date and calendar calculations with explicit conventions.
- Table, spreadsheet, and structured-data calculations.
- Deterministic simulations, data transformations, and lightweight statistical analysis.
- Algorithmic tasks or repetitive procedures with clearly defined rules.
A simple calculator call or fixed function may be a better choice for a single operation. PAL becomes more useful when the task needs multiple linked operations, branching logic, or state that is awkward to express as one narrow tool call.
What PAL cannot fix
Execution makes the code’s behavior more reproducible; it does not prove the code represents the question correctly. A Python interpreter can execute incorrect logic perfectly. PAL shifts some risk away from arithmetic and toward interpretation, specification, and program design.
- Misread questions or extracted values: If the model reads 15% as 15 rather than 0.15, execution will faithfully calculate from the wrong value.
- Wrong formulas or assumptions: A valid program can apply the wrong tax order, unit conversion, date convention, or business rule.
- Ambiguity and missing facts: Code cannot resolve requirements or factual inputs that were never established.
- Subjective tasks: Tone, cultural context, judgment, and incomplete evidence are not made deterministic by writing code.
- Runtime and data problems: Bugs, stale files, poor-quality source data, or incorrect external tool results can still produce bad answers.
Floating-point arithmetic is another edge case: binary floating-point can represent some decimal fractions imprecisely. For financial work, use decimal arithmetic or integer minor units, and show the units in the output. Date, indexing, and counting problems should be tested around boundary cases to catch off-by-one errors.
When is an executable answer trustworthy?
“The code ran” is only one check. Evaluate a PAL result on separate levels:
- Syntactic validity: Did the program parse and run?
- Execution correctness: Did the runtime produce the output implied by that code?
- Semantic correctness: Does the program model the original question and its rules?
- Factual correctness: Are the input values and assumptions true?
- Safety: Was executing the program harmless and adequately contained?
For consequential calculations, ask the system to report assumptions, program, result, and validation separately. Independently check critical outputs with a second calculation, a domain rule, or a deterministic test. Code is evidence of a procedure, not proof of truth.
How should a PAL implementation handle errors?
A production workflow needs controls around generation and execution, not just a call to Python. A bounded recovery sequence can look like this:
Recommended Free Tools
- Generate a program from the question.
- Apply static checks and reject disallowed operations before execution.
- Run it in an isolated environment with time and resource limits.
- Capture the exit status, standard output, standard error, and resource use as typed results.
- If execution fails, return the error details to the model for a limited repair attempt.
- Validate a successful result independently; return it, or escalate when confidence or impact warrants human review.
Do not allow unrestricted self-repair. Repeated retries can add cost, hide uncertainty, or change a sound approach into a faulty one. Set a maximum repair count and define when the system must stop and ask for human help.
How can you build a safer PAL-style prototype?
A minimal computational function might look like this:
def solve():
items = [12, 15, 8]
subtotal = sum(items)
tax = subtotal * 0.08
return round(subtotal + tax, 2)
print(solve())
The arithmetic is simple; the hard production decision is where and how to run generated code. Treat model-generated code as untrusted input. Never execute it directly on the application host.
- Use an isolated container, restricted subprocess, WebAssembly runtime, or managed execution service.
- Run under a non-privileged user; disable network access unless a specific, reviewed requirement needs it.
- Do not mount sensitive host directories or expose environment variables and credentials.
- Set hard limits for CPU, memory, processes, file size, and wall-clock time.
- Restrict imports and inspect code before execution; for narrow tasks, consider a fixed function interface or domain-specific language instead of unrestricted Python.
- Capture execution metadata and logs, while controlling what prompt or data may be retained.
- Protect code-generation flows that read documents or spreadsheets from prompt injection: treat their contents as untrusted data, not instructions.
How does PAL relate to current hosted code-execution tools?
Modern hosted tools can provide execution as one component of a larger model API. They are not necessarily implementations of the original PAL method, and support depends on the provider, model, API surface, account, and configuration.
Rank #4
Gemini API
Google documents code execution for generating and running code, including workflows with text and CSV files and graph output. Its documented Gemini code-execution environment has a maximum runtime of 30 seconds; that limit should not be generalized to every Google AI product. The documentation says code can be regenerated after an error up to five times. See Gemini code-execution documentation.
Google says enabling code execution has no separate charge; paid API use still incurs the model’s standard token charges. The pricing page also describes Google AI Studio as free in available regions, subject to limits and changes. Pricing and availability can change, so check the current terms for your region and account. See Gemini API pricing.
OpenAI API
OpenAI’s model documentation lists code interpreter among the tools supported for GPT-5.4, alongside hosted shell and other capabilities; the available configuration depends on the API surface, model, account, and tool setup. See the GPT-5.4 documentation. OpenAI also documents usage reporting that includes code-interpreter sessions. See the usage API reference.
An earlier Responses API announcement listed a $0.03-per-container Code Interpreter charge. That is a historical price signal, not a confirmed universal current rate; check current product-specific billing before estimating costs. See the announcement.
Amazon Bedrock
Bedrock offers model access and enterprise cloud controls, but it is not by itself a complete PAL framework: an implementation still needs orchestration, execution, validation, and safety decisions. AWS pricing varies by provider, model, region, tier, and inference mode; selected models may be available for batch inference at a 50% discount versus on-demand pricing. See Bedrock pricing.
Best Value
AWS announced general availability of OpenAI models and Codex on Bedrock on June 1, 2026, with pay-per-token pricing and usage counting toward existing AWS commitments. Availability and terms should be checked for the intended account and region. See the AWS announcement.
Self-managed execution
A self-hosted stack can combine a locally served or open-weight LLM, an orchestrator, a sandbox, logging, and an evaluation harness. It offers greater control over data flow and runtime configuration, but the organization must also maintain the infrastructure, model serving, security, and reliability. The original PAL repository is an Apache-2.0 reference implementation showing an LLM backend connected to a Python backend; its examples use historical model and API identifiers, so they are not current production setup instructions. See the PAL repository.
| Option | Main purchase or cost | Strongest advantage | Main drawback |
|---|---|---|---|
| OpenAI API tools | Model/API usage; tool billing depends on configuration | Integrated model and tool ecosystem | Vendor dependence and configuration-specific pricing |
| Gemini API Code Execution | Model/API usage; no separate charge to enable code execution per Google | Managed execution with documented file and graph workflows | Documented runtime and capability limits |
| Amazon Bedrock | Provider-, model-, region-, and tier-dependent cloud usage | Enterprise governance and provider choice | More platform complexity; PAL orchestration remains an implementation task |
| Self-hosted stack | Infrastructure, engineering, and operations | Control, privacy, and customization | Highest operational burden |
How should you evaluate a PAL system?
Compare the PAL workflow with a non-executing baseline on the same representative tasks. Measure more than whether the final answer is right:
- Exact-answer accuracy and semantic correctness.
- Program execution success rate and frequency of repair attempts.
- Performance on ambiguous, malformed, and boundary-case inputs.
- Latency, model-token usage, and runtime cost.
- Validation, abstention, and escalation quality.
- Security outcomes, including blocked unsafe operations and data exposure risks.
Keep separate scores for code that runs and code that correctly represents the task. Otherwise, a high execution-success rate can disguise a system that reliably executes the wrong procedure.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

