Sometimes. Language models can apply patterns to genuinely new examples, particularly when a task shows how familiar pieces fit together. But success depends on the rule, the examples, and what is held out for testing. A correct answer alone cannot tell us whether a model learned a general rule, reused a skill it already had, or arrived at the answer another way.
Start with a pattern—and ask what would count as learning it
Suppose a prompt gives two examples:
red bluebecomesblue red.red red bluebecomesblue red red.
A natural guess is “reverse the order.” But these examples are also compatible with narrower rules, such as moving the last item to the front. Both rules predict the same answers so far. A model that gets another example right has not necessarily shown which rule it inferred; the test needs a case where the competing rules make different predictions.
This is the central difficulty in judging rule learning: examples can support more than one explanation. The useful question is not simply whether a model answers correctly, but whether it succeeds on cases that distinguish the intended rule from plausible alternatives.
What kind of generalization is being tested?
“Generalizing to a new example” can mean several different things. Researchers use compositional generalization for handling a new combination of components the model has encountered before. In-context learning means responding to examples supplied in a prompt, without fine-tuning the model for that task. Neither term, by itself, establishes that the model has formed a symbolic rule like a person might describe in words.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →| Test condition | What is new? | What success can show |
|---|---|---|
| Unseen combination | Known components appear in a combination not shown in the demonstrations. | The model may be combining familiar skills systematically; success does not establish how it represents the rule. |
| Novel word or symbol | The relevant structure is expressed with unfamiliar labels rather than familiar words. | Success is less easily explained by familiarity with those particular words, though it still does not prove human-like understanding. |
| Longer sequence or new structure | The example extends a pattern or changes its structure beyond the demonstrated cases. | This probes whether performance transfers beyond short or structurally similar examples. |
| Rule-violating prompt | At least one rule in a formal-language prompt is violated. | This tests extrapolation under a specifically defined out-of-distribution change, rather than ordinary completion of familiar examples. |
These conditions are not interchangeable. A strong result on a new combination of known words does not automatically predict success with unfamiliar symbols, longer sequences, or a prompt that breaks a formal rule. The NeurIPS study by Mészáros and colleagues calls the last kind of out-of-distribution case “rule extrapolation” and examines it in formal languages; its definition makes clear why an evaluation must specify what changed between demonstrations and test prompts (NeurIPS Proceedings, 2024).
When language models do generalize
There is evidence for rule-like generalization under particular conditions. Song, Xu, and Zhong study hidden-rule and symbolic reasoning tasks, reporting that compositional structure matters for out-of-distribution generalization in the settings they examine. They also caution that the mechanisms behind such generalization remain poorly understood (PNAS, 2025).
Rank #2
Chen and colleagues report that a prompt format called Skills-in-Context can elicit systematic generalization in their tested tasks by demonstrating foundational skills as well as examples that compose those skills. Their paper reports near-perfect results for those tasks with as few as two exemplars. That is a result for the evaluated tasks and prompt method—not evidence that two examples will teach any model any rule, or that the model discovered a universal procedure (Findings of EMNLP, 2024).
Prompt examples themselves make a difference. In experiments on in-context compositional generalization, An and colleagues find better results when demonstrations are structurally similar to the test case, diverse from one another, and individually simple. They also report weaker generalization with fictional words and emphasize covering the linguistic structures needed for the test. This means apparent rule use can depend partly on whether the prompt provides useful structural coverage and whether the model is already familiar with the language it sees (ACL, 2023).
Free tools Windows power users keep installed
One-click scans. No signup required.
Rank #3
Strong benchmark scores do not mean generalization is solved
Performance can vary sharply across ways of testing systematicity. Lake and Baroni’s meta-learning compositional model achieved at least 99.78% accuracy on three SCAN systematic-generalization splits involving lexical generalization. In the same study, it failed on other structural generalization tasks. The result is therefore evidence of excellent performance on those splits, not a general score for rule learning. As the authors put it, “Systematicity continues to challenge models” (Nature, 2023).
A separate study by Hosseini and colleagues reports that the relative compositional generalization gap decreased as model scale increased across four model families and three semantic-parsing datasets. That trend applies to those evaluations; it does not show that scaling eliminates compositional limits across tasks or model families (BlackboxNLP, 2022).
Rank #4
- Logic Puzzles for Kids Ages 4-8
- Brand : Spotlight Media
Taken together, these findings argue against both simple extremes: that language models only copy examples, and that they reliably infer the intended rule behind any pattern. Researchers have observed generalization, but it is conditional and uneven. The papers do not establish a single population-wide rate at which models learn rules.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.How to tell whether a model learned more than the examples
If you are testing a model on a pattern task, design the evaluation so that a memorized or overly narrow answer strategy cannot pass by accident:
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Best Value
- State the intended rule precisely. List the parts of an example that matter and what transformation they undergo. If two rules fit the demonstrations, make that ambiguity explicit rather than treating one answer as proof of the intended interpretation.
- Separate demonstrations from test cases. Hold back examples until evaluation, and say what is withheld: a combination of known parts, a new symbol, a longer sequence, or a different structure. A test that resembles a demonstration in every important way provides weaker evidence of transfer than one that changes a specified feature.
- Use contrastive cases. Include examples where competing explanations predict different outputs. Check both cases the rule permits and cases it excludes; otherwise a model may succeed with a shortcut that matches only the positive examples.
- Control familiarity and prompt coverage. Try familiar language and unfamiliar symbols separately, and make sure demonstrations cover the structures needed for the test. Vary the examples’ similarity, diversity, and complexity when comparing prompts, because those properties can affect in-context results.
- Report the boundary of the result. Give the tested task, model or method, demonstrations, held-out condition, and accuracy together. A benchmark score supports a claim about that setup—not about all patterns or about the model’s internal reasoning.
Even a carefully passed test demonstrates behavior, not mechanism. Output alone cannot settle whether a model represents an explicit symbolic rule, combines previously learned skills, or uses another learned process. Song and colleagues note that mechanisms of out-of-distribution generalization remain poorly understood, so claims about internal rule use should be narrower than claims about observed performance.
So, can an LLM apply a rule to a new example?
Yes, sometimes: experiments show models and model-based systems succeeding on specified forms of compositional and out-of-distribution generalization. But results change with the task, demonstrations, symbols, and held-out structure. A convincing claim that a model learned a rule therefore needs more than a correct familiar-looking answer: it needs a test designed to distinguish genuine transfer from simpler ways of matching the examples.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




