October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsClean PCRecommendedOne scan can reveal what keeps slowing WindowsLook for cleanup and repair opportunities.Run ScanOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
World desk5 min

Can a Language Model Learn the Rule Behind a Pattern?

Language models sometimes apply patterns to unseen cases, especially when prompts show how familiar skills combine. But performance depends on the task and test, and a correct answer does not prove a general rule was learned.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Sometimes. Language models can apply patterns to genuinely new examples, particularly when a task shows how familiar pieces fit together. But success depends on the rule, the examples, and what is held out for testing. A correct answer alone cannot tell us whether a model learned a general rule, reused a skill it already had, or arrived at the answer another way.

Start with a pattern—and ask what would count as learning it

Suppose a prompt gives two examples:

  • red blue becomes blue red.
  • red red blue becomes blue red red.

A natural guess is “reverse the order.” But these examples are also compatible with narrower rules, such as moving the last item to the front. Both rules predict the same answers so far. A model that gets another example right has not necessarily shown which rule it inferred; the test needs a case where the competing rules make different predictions.

This is the central difficulty in judging rule learning: examples can support more than one explanation. The useful question is not simply whether a model answers correctly, but whether it succeeds on cases that distinguish the intended rule from plausible alternatives.

What kind of generalization is being tested?

“Generalizing to a new example” can mean several different things. Researchers use compositional generalization for handling a new combination of components the model has encountered before. In-context learning means responding to examples supplied in a prompt, without fine-tuning the model for that task. Neither term, by itself, establishes that the model has formed a symbolic rule like a person might describe in words.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Test condition What is new? What success can show
Unseen combination Known components appear in a combination not shown in the demonstrations. The model may be combining familiar skills systematically; success does not establish how it represents the rule.
Novel word or symbol The relevant structure is expressed with unfamiliar labels rather than familiar words. Success is less easily explained by familiarity with those particular words, though it still does not prove human-like understanding.
Longer sequence or new structure The example extends a pattern or changes its structure beyond the demonstrated cases. This probes whether performance transfers beyond short or structurally similar examples.
Rule-violating prompt At least one rule in a formal-language prompt is violated. This tests extrapolation under a specifically defined out-of-distribution change, rather than ordinary completion of familiar examples.

These conditions are not interchangeable. A strong result on a new combination of known words does not automatically predict success with unfamiliar symbols, longer sequences, or a prompt that breaks a formal rule. The NeurIPS study by Mészáros and colleagues calls the last kind of out-of-distribution case “rule extrapolation” and examines it in formal languages; its definition makes clear why an evaluation must specify what changed between demonstrations and test prompts (NeurIPS Proceedings, 2024).

When language models do generalize

There is evidence for rule-like generalization under particular conditions. Song, Xu, and Zhong study hidden-rule and symbolic reasoning tasks, reporting that compositional structure matters for out-of-distribution generalization in the settings they examine. They also caution that the mechanisms behind such generalization remain poorly understood (PNAS, 2025).

Chen and colleagues report that a prompt format called Skills-in-Context can elicit systematic generalization in their tested tasks by demonstrating foundational skills as well as examples that compose those skills. Their paper reports near-perfect results for those tasks with as few as two exemplars. That is a result for the evaluated tasks and prompt method—not evidence that two examples will teach any model any rule, or that the model discovered a universal procedure (Findings of EMNLP, 2024).

Prompt examples themselves make a difference. In experiments on in-context compositional generalization, An and colleagues find better results when demonstrations are structurally similar to the test case, diverse from one another, and individually simple. They also report weaker generalization with fictional words and emphasize covering the linguistic structures needed for the test. This means apparent rule use can depend partly on whether the prompt provides useful structural coverage and whether the model is already familiar with the language it sees (ACL, 2023).

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Strong benchmark scores do not mean generalization is solved

Performance can vary sharply across ways of testing systematicity. Lake and Baroni’s meta-learning compositional model achieved at least 99.78% accuracy on three SCAN systematic-generalization splits involving lexical generalization. In the same study, it failed on other structural generalization tasks. The result is therefore evidence of excellent performance on those splits, not a general score for rule learning. As the authors put it, “Systematicity continues to challenge models” (Nature, 2023).

A separate study by Hosseini and colleagues reports that the relative compositional generalization gap decreased as model scale increased across four model families and three semantic-parsing datasets. That trend applies to those evaluations; it does not show that scaling eliminates compositional limits across tasks or model families (BlackboxNLP, 2022).

Taken together, these findings argue against both simple extremes: that language models only copy examples, and that they reliably infer the intended rule behind any pattern. Researchers have observed generalization, but it is conditional and uneven. The papers do not establish a single population-wide rate at which models learn rules.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to tell whether a model learned more than the examples

If you are testing a model on a pattern task, design the evaluation so that a memorized or overly narrow answer strategy cannot pass by accident:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. State the intended rule precisely. List the parts of an example that matter and what transformation they undergo. If two rules fit the demonstrations, make that ambiguity explicit rather than treating one answer as proof of the intended interpretation.
  2. Separate demonstrations from test cases. Hold back examples until evaluation, and say what is withheld: a combination of known parts, a new symbol, a longer sequence, or a different structure. A test that resembles a demonstration in every important way provides weaker evidence of transfer than one that changes a specified feature.
  3. Use contrastive cases. Include examples where competing explanations predict different outputs. Check both cases the rule permits and cases it excludes; otherwise a model may succeed with a shortcut that matches only the positive examples.
  4. Control familiarity and prompt coverage. Try familiar language and unfamiliar symbols separately, and make sure demonstrations cover the structures needed for the test. Vary the examples’ similarity, diversity, and complexity when comparing prompts, because those properties can affect in-context results.
  5. Report the boundary of the result. Give the tested task, model or method, demonstrations, held-out condition, and accuracy together. A benchmark score supports a claim about that setup—not about all patterns or about the model’s internal reasoning.

Even a carefully passed test demonstrates behavior, not mechanism. Output alone cannot settle whether a model represents an explicit symbolic rule, combines previously learned skills, or uses another learned process. Song and colleagues note that mechanisms of out-of-distribution generalization remain poorly understood, so claims about internal rule use should be narrower than claims about observed performance.

So, can an LLM apply a rule to a new example?

Yes, sometimes: experiments show models and model-based systems succeeding on specified forms of compositional and out-of-distribution generalization. But results change with the task, demonstrations, symbols, and held-out structure. A convincing claim that a model learned a rule therefore needs more than a correct familiar-looking answer: it needs a test designed to distinguish genuine transfer from simpler ways of matching the examples.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
Crashes, No Sound, or Screen Glitches?Free driver scan
Windows Errors? Fix Them Before They SpreadFree repair scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.