Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

LLM coding tools do not reliably make experienced developers faster in every setting. A 2025 randomized trial found that experienced contributors to mature open-source projects took 19% longer on realistic tasks when using early-2025 AI tools. A 2026 follow-up produced results consistent with a speedup, but its researchers said selection and measurement problems made the size of any gain unreliable. The most defensible answer is conditional: impact depends on the task, repository, tool and workflow—and must be measured against quality and downstream cost, not code volume alone.

First decide what “productivity” means

A developer can type less, finish a ticket sooner, ship more features or create more value. Those are related, but they are not interchangeable:

  • Speed: time to complete a defined task. This is useful for comparable bugs, features or refactors, but a fast first patch may still need extensive review or repair.
  • Output: accepted work delivered over a period—such as merged changes, resolved incidents or releases. More output is not necessarily more valuable.
  • Value: user or business outcomes, such as improved reliability, lower infrastructure costs or customer adoption. These may take time to emerge and can be hard to attribute to one change.
  • Sustainable engineering capacity: the amount of useful, reliable software a team can deliver without accumulating defects, rework, security exposure, maintenance burden or burnout.

For an engineering organization, sustainable capacity is usually the most useful goal. METR’s 2026 survey work also distinguishes speed from value: AI may make a given task faster, enable work that otherwise would not be attempted, or do both. A task that would never have entered the backlog cannot show up as a faster ticket in a same-task experiment.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A practical accounting is: net productivity gain = generation time saved − context, verification, correction, integration and maintenance costs. Less typing is only one part of that equation.

#1 Best Overall
Kisnt KN85 Wireless Mechanical Keyboard, 75% Layout, Bluetooth/2.4GHz/USB-C, Custom RGB Backlit, Hot-Swappable Linear Switch, Creamy Sound for Gaming/Typing (Retro Beige)
  • 【75% Space‑saving Layout】The KN85 series is a compact 85‑key keyboard (13.68" × 5.51" × 1.77") that keeps all the essentials (F1–F12, arrows, shortcuts) without the number pad. It frees up 25% of desk space for better mouse movement. Designed for small desks, laptop setups, gamers and minimalists. For frequent number‑pad input, choose our full‑size KN104 with a complete dedicated numpad, or opt for our new KN98 model — compact 99‑key that retains the numpad while saving desktop real‑estate
  • 【Tri-Mode Connectivity for Multi-Device Workflow】Connect via USB‑C, 2.4GHz wireless, or Bluetooth 5.0 (3 channels supported), with ultra‑low latency (USB 2ms, 2.4G 5ms, BT 11ms). Switch seamlessly between Windows and Mac to work across your PC, laptop, tablet, smartphone, or gaming console. Perfect for programmer, student, creator, or hybrid worker. The built‑in 4000mAh rechargeable battery ensures stable wireless performance. Continue typing while charging via wired mode when power runs low
  • 【Creamy Thocky Typing Sound】The gasket mount absorbs harsh vibrations and hollow echoes to produce a smooth marbly thock, rather than loud clacky taps. Each keypress feels softly cushioned. Whether you’re working late at home or typing in a shared office space, the mellow, ASMR-like tone makes every keystroke a genuinely enjoyable experience
  • 【Hot-swap for Tailored Sound & Tactile】Pre-lubed Bsun linear switches (45-50gf actuation) deliver a buttery response. Compatible with both 3 pin and 5pin switches, they enables solder-free swapping. From beginners to frequent typists and dedicated writers, craft your preferred typing signature without complex modding
  • 【RGB Backlighting & Programmable】A warm ambient glow surrounds PBT keycaps and case edges, creating a calm, inviting desk vibe for late-night workspace. Adjust hues and brightness through shortcut keys or companion software. The KN85 driver (Windows only, wired/2.4G mode) lets you remap keys and set custom macros to boost your daily productivity

What the 2025 METR experiment actually found

In a randomized controlled trial published July 10, 2025, METR studied 16 experienced open-source developers completing 246 real issues in repositories they had contributed to for years. The repositories averaged more than 22,000 stars and one million lines of code. Tasks included features, bug fixes and refactors, and took roughly two hours on average. Participants were assigned to work with AI allowed or disallowed; the AI condition primarily used Cursor Pro with Claude 3.5 or 3.7 Sonnet, then-current frontier tools. METR’s study and methods describe the setting and results.

The measured result was a 19% increase in task-completion time when AI was allowed, with a confidence interval consistent with roughly 2% to 39% more time. Before the tasks, developers expected AI to make them about 24% faster; afterward, they still estimated a speedup of about 20%, despite the measured slowdown.

This is important evidence, not a universal rule. It applies to a small group of experienced contributors, familiar and mature repositories, a particular set of realistic tasks, early-2025 tools and the study’s workflow and success standards. It does not establish that AI slows most developers, or that the same result applies to beginners, greenfield projects, unfamiliar codebases or later tools. METR itself cautioned against those broader conclusions.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The trial matters partly because it measured work in existing repositories rather than whether a model could solve a self-contained benchmark problem. A plausible patch still has to satisfy the project’s tests, style, documentation, compatibility and review expectations. Producing code quickly can therefore coexist with taking longer to finish acceptable work.

Rank #2
Sale
AULA F75 Pro Wireless Mechanical Keyboard,75% Hot Swappable Custom Keyboard with Knob,RGB Backlit,Pre-lubed Reaper Switches,Side Printed PBT Keycaps,2.4GHz/USB-C/BT5.0 Mechanical Gaming Keyboards
  • Tri-mode Connection Keyboard: AULA F75 Pro wireless mechanical keyboards work with Bluetooth 5.0, 2.4GHz wireless and USB wired connection, can connect up to five devices at the same time, and easily switch by shortcut keys or side button. F75 Pro computer keyboard is suitable for PC, laptops, tablets, mobile phones, PS, XBOX etc, to meet all the needs of users. In addition, the rechargeable keyboard is equipped with a 4000mAh large-capacity battery, which has long-lasting battery life
  • Hot-swap Custom Keyboard: This custom mechanical keyboard with hot-swappable base supports 3-pin or 5-pin switches replacement. Even keyboard beginners can easily DIY there own keyboards without soldering issue. F75 Pro gaming keyboards equipped with pre-lubricated stabilizers and LEOBOG reaper switches, bring smooth typing feeling and pleasant creamy mechanical sound, provide fast response for exciting game
  • Advanced Structure and PCB Single Key Slotting: This thocky heavy mechanical keyboard features a advanced structure, extended integrated silicone pad, and PCB single key slotting, better optimizes resilience and stability, making the hand feel softer and more elastic. Five layers of filling silencer fills the gap between the PCB, the positioning plate and the shaft,effectively counteracting the cavity noise sound of the shaft hitting the positioning plate, and providing a solid feel
  • 16.8 Million RGB Backlit: F75 Pro light up led keyboard features 16.8 million RGB lighting color. With 16 pre-set lighting effects to add a great atmosphere to the game. And supports 10 cool music rhythm lighting effects with driver. Lighting brightness and speed can be adjusted by the knob or the FN + key combination. You can select the single color effect as wish. And you can turn off the backlight if you do not need it
  • Professional Gaming Keyboard: No matter the outlook, the construction, or the function, F75 Pro mechanical keyboard is definitely a professional gaming keyboard. This 81-key 75% layout compact keyboard can save more desktop space while retaining the necessary arrow keys for gaming. Additionally, with the multi-function knob, you can easily control the backlight and Media. Keys macro programmable, you can customize the function of single key or key combination function through F75 driver to increase the probability of winning the game and improve the work efficiency. N key rollover, and supports WIN key lock to prevent accidental touches in intense games

Why AI can slow a developer who already knows the code

The study’s result does not mean any single mechanism explains the slowdown. METR examined potential explanations and reported evidence implicating several factors; the following are mechanisms that help explain why generation speed need not become end-to-end speed.

  • Repository context is expensive. A model may not know the historical reason for an abstraction, a maintainer’s unwritten convention or a dependency across modules. The developer has to supply context, identify what the model missed and repair the result.
  • Verification does not disappear. Generated changes still need tests, review, type checks, linting, security inspection, compatibility checks and sometimes performance analysis. In a high-standard codebase, checking an unfamiliar patch can cost more than writing a small change directly.
  • Familiarity can make direct work quicker. A maintainer may already know where a fix belongs and which edge cases matter. Describing the work, waiting for a response, reviewing it and correcting a wrong-file or wrong-assumption edit can add steps.
  • Interaction has overhead. Prompt construction, retries, waiting, context switching among editor, terminal and chat, and reverting unwanted edits all consume time.
  • Expectations can be miscalibrated. In this particular trial, participants expected and perceived a speedup that the task-time measure did not show. That gap is a reason to measure actual workflow outcomes, not evidence that every developer misjudges every tool.

METR also noted that its tested setup may not represent the best possible prompting, sampling or scaffolding. The result is not a ceiling on what better tools or workflows might achieve.

Why benchmarks, surveys and real-world trials can disagree

These evidence types answer different questions. Treating them as competing votes on one universal productivity number creates confusion.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Evidence What it can tell you What it may miss
Controlled human experiment Whether a specified AI workflow changed outcomes for people doing assigned work under study conditions. Small samples, tool-version drift, selection effects and limited representation of long-term or changing workflows.
Agent coding benchmark Whether a system can solve a set of defined coding tasks under a specified scaffold and scoring rule. Human prompting and review time, repository-specific conventions, maintainability and operational requirements. Automated pass/fail scoring is not the same as a maintainer accepting a production change.
Survey or anecdote Perceived usefulness, adoption, task expansion, frustration and changes in how people work. A reliable counterfactual. People may report typing speed as total time saved, or remember successful cases more readily than expensive failures.
Production telemetry How delivery, review, quality and operations changed in a real organization over time. Causation: staffing, task mix, release policy, seasonality and other changes can move alongside AI adoption.

METR explicitly contrasts its real-task trial and human acceptability criteria with coding benchmarks such as SWE-Bench Verified and RE-Bench, which use algorithmic scoring and often more autonomous scaffolding. Benchmark performance can indicate model capability; it is not, by itself, a measured productivity gain for an experienced developer.

Rank #3
Sale
Logitech MX Keys S Wireless Keyboard Low Profile Fluid Precise - Graphite
  • Fluid Typing Experience: Laptop-like profile with spherically-dished keys shaped for your fingertips delivers a fast, fluid, precise and quieter typing experience
  • Automate Repetitive Tasks: Easily create and share time-saving Smart Actions shortcuts to perform multiple actions with a single keystroke with the Logi Options+ app (1)
  • Smarter Illumination: Backlit keyboard keys light up as your hands approach and adapt to the environment; Now with more lighting customizations on Logi Options+ (1)
  • More Comfort, Deeper Focus: Work for longer with a solid build, low-profile design and an optimum keyboard angle that is better for your wrist posture
  • Multi-Device, Multi OS Bluetooth Keyboard: Pair with up to 3 devices on nearly any operating system (Windows, macOS, Linux) via Bluetooth Low Energy or included Logi Bolt USB receiver (2)

What changed in the 2026 follow-up

METR’s February 2026 follow-up broadened the study to 57 developers, 143 repositories and more than 800 tasks, including 10 participants from the earlier experiment. The participant pool had a median of 10 years’ experience and included smaller, greener and less mature repositories. Raw estimates suggested an 18% speedup for returning participants (with an interval ranging from 38% faster to 9% slower) and a 4% speedup for newly recruited developers (interval from 15% faster to 9% slower). METR’s update explains why it did not treat these figures as a reliable estimate of current uplift.

The central problem was selection. As AI use became more common, some developers were unwilling to participate if they might have to work without it. In surveys, 30%–50% said they had avoided submitting some tasks because they did not want those tasks assigned to an AI-disallowed condition. That can change who enters a study and which work is available to assign, making the comparison less representative.

Time attribution also became harder. Some participants ran multiple agents concurrently or did other work while agents waited, making “time spent on the task” less clear. METR’s conclusion was that developers were probably more accelerated by AI in early 2026 than in early 2025, but that the follow-up supported the size of that increase only weakly. It is evidence that the 2025 slowdown may not describe newer workflows; it is not a clean reversal with a dependable percentage gain.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What the 2026 survey adds—and what it cannot prove

A METR survey conducted February–April 2026 included 349 technical workers, of whom 87 were software engineers. Respondents reported median AI-related changes in the value of their work of roughly 1.4× to 2×, and a median speed change of 3×. METR warned that speed reports likely overstate value gains. Participants retrospectively estimated a 1.3× change in work value in March 2025 and 2× in March 2026, and forecast 2.5× in March 2027. These are reported perceptions and forecasts, not controlled measurements of delivered output. See the survey results and METR’s discussion of task substitution and uplift.

Rank #4
Sale
AULA F99 Wireless Mechanical Keyboard,Tri-Mode BT5.0/2.4GHz/USB-C Hot Swappable Custom Keyboard,Pre-lubed Linear Switches,RGB Backlit Computer Gaming Keyboards for PC/Tablet/PS/Xbox
  • Multi-Device Connection: The F99 wireless mechanical keyboard provides three connection methods, including BT5.0, 2.4GHz wireless mode, and USB wired mode. It can be connected to up to five devices at the same time, and switch between them easily by FN and key combination keys. No limits about your keyboard connection to meet the needs of work, gaming, and study
  • Hot-swappable Custom Keyboard: The switches and keycaps can be freely replaced(keycap/switch puller are included in the package).This customizable keyboard with hot-swap PCB allows users to replace 3 pins/5 pins switches easily without soldering issue. F99 mechanical keyboards equipped with pre-lubed linear switches, bring smooth typing feeling and pleasant typing sound, provide fast response for exciting game
  • Mechanical Gaming Keyboard: F99 is a premium mechanical keyboard for both work and game. With 16 RGB lighting effect to adds a great atmosphere to the game room. Keys support macro customization, which allows macro recording and editing, customize key function and 16.8 million light colors, and supports cool music rhythm lighting effects with driver. N-key rollover, keyboard can respond to multiple key presses at the same time, which is helpful in very exciting real-time games
  • Gasket Structure and PCB Single Key Slotting: This computer keyboard features a advanced structure, extended integrated silicone pad, and PCB single key slotting, better optimizes resilience and stability, making the hand feel softer and more elastic. Five layers of filling silencer fills the gap between the PCB, the positioning plate and the shaft,effectively counteracting the cavity noise sound of the shaft hitting the positioning plate, and providing a solid feel
  • PBT Keycaps and 8000mAh Battery: 99 keys 96% layout compact keyboard can save more desktop space while keep necessary arrow keys and number area for games and work. The rechargeable keyboard built-in 8000mAh large capcacity battery to provide more power and longer battery life. Double shot PBT keycaps, made from two colors material molded into each others, make the keycaps characters maintain the vibrance and saturation, clear and not fade

The survey is useful for understanding how workers perceive AI and how it may change the work they choose to do. It cannot establish that a team delivered twice the value, or that an individual completed equivalent tasks twice as fast. A developer might use AI to add tests, automate a migration or prototype an idea that would otherwise be deferred. That is potential value expansion, not necessarily acceleration on a preselected task.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

A practical way to measure AI’s impact in your organization

Do not run a trial called simply “AI versus no AI.” Define the workflow, decide what success means, and track its costs as well as its benefits.

  1. Specify the treatment. Record the product, model and version; whether autocomplete, chat, editor or terminal agents are included; whether web search is allowed; whether multiple agents may run; permitted shell and network access; and whether AI may be used for tests, documentation and debugging. Include training and usage logging. “AI allowed” can mean very different things to different developers.
  2. Segment the work. At minimum, classify bugs, features, refactors and greenfield tasks; familiar versus unfamiliar subsystems; small versus cross-repository changes; test coverage; tacit knowledge; risk level; and human-led versus agent-led work. Aggregate averages can hide gains on boilerplate and losses on subtle changes.
  3. Establish a baseline and comparable tasks. Use historical data cautiously, because teams and tools change. Where feasible, randomly assign comparable tasks or developers to workflows. A task-level randomized trial can estimate effects on similar work, but may become unrealistic if developers selectively submit AI-friendly tasks or refuse the control condition. A developer- or team-level trial captures sustained workflow better, but needs more participants and is more exposed to differences between teams.
  4. Measure quality-adjusted completion. Track time through review and acceptance, not just time to first patch or passing test. Include rework, requested changes, failed tests, follow-up fixes, reverts and post-release defects. Compare work that meets the same acceptance and quality bar.
  5. Track several kinds of outcome. Combine delivery and cycle time with review latency, escaped defects, change failures, incident recovery, maintenance burden and relevant business outcomes. Use medians and distributions as well as averages: a tool may create many quick wins and a few costly failures.
  6. Ask developers separately. Survey perceived usefulness, cognitive load, interruptions, trust calibration, learning and frustration. These are meaningful outcomes, but label them as perceptions rather than measured time savings. Interviews and task transcripts can reveal where context or verification costs arise.
  7. Log versions and concurrency. Models and products change; record the exact configuration and dates. For concurrent agents, record wall-clock time, human attention and agent activity separately where possible. Do not assume elapsed time, active developer time and total compute time are the same measure.
  8. Repeat the evaluation. Reassess after major model, product, training or workflow changes. Report results by task type and developer experience, not just as a single company-wide percentage.

Metrics that mislead when used alone

  • Lines of code: more code may mean duplication or future maintenance, not more value.
  • Pull-request count: a higher count can reflect trivial, fragmented or generated changes.
  • AI-generated-code share: measures tool use, not correctness, accepted work or economic benefit.
  • Self-reported time saved: useful for adoption research, but not a replacement for observed outcomes. The 2025 METR trial found a notable gap between expected/perceived speedup and measured task time.
  • Benchmark scores: measure performance on specified benchmark tasks, not the full human workflow in your repositories.
  • Token or agent-call volume: measures consumption, not productive output.

Useful delivery indicators may include task-completion distributions, pull-request cycle time, lead time, deployment frequency and incident recovery time. Pair them with quality indicators such as defects before and after release, reverts, hotfixes, test failures, review requests, security findings and rework. Maintainability measures can include follow-up fixes, complexity, duplication, documentation and the time a developer later spends understanding the change. Choose metrics appropriate to the work; no single dashboard can turn unlike tasks into a fair comparison.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Where AI is more or less likely to fit

AI assistance is a stronger candidate for broad use when work is repetitive and well-scoped, the code has good tests and documentation, developers can check results quickly, and usage fits security and privacy requirements. Examples to evaluate include boilerplate, test scaffolding, documentation drafts, routine API migrations and exploration of unfamiliar libraries.

Best Value
Sale
Redragon K668 108-Key Hot-Swap Wired RGB Gaming Keyboard, Extra 4 Hotkeys
  • 4 Extra Hotkeys, Full-Size 108-Key Anti-Ghosting - Dedicated shortcut keys default to mute, calculator, screen lock and desktop, while 104 keys register accurately even during rapid multi-key combos.
  • Swap Switches Without Soldering, Smooth and Quiet - The upgraded socket accepts almost any 3-pin or 5-pin switch, and stock Red linear switches keep clicks discreet for shared spaces.
  • Vibrant RGB for a True eSports Vibe - Up to 19 preset lighting modes with adjustable brightness and flow speed, including a music-sync mode that lights up in time with your desktop audio.
  • Ergonomic 2-Stage Feet, 2 Sets of Mixed Color Keycaps - Adjustable feet relax your wrists during long sessions, and two included keycap sets let you swap looks whenever you want a fresh vibe.
  • Pro Software for Even Deeper Customization - Reassign the 4 hotkeys to your own shortcuts, design custom lighting effects, and program macros with your own keybindings.

Use tighter controls or selective assistance when requirements are ambiguous, repository knowledge is largely tacit, tests are weak, architecture is fragile, or a plausible but incorrect change has high security or operational cost. These conditions do not prove AI will hurt; they raise the cost of verification and make measurement especially important.

Agentic workflows are most plausible when work can be decomposed, execution is sandboxed, tests can run automatically, human permissions are explicit and asynchronous work offers a real advantage. They can also add coordination, merge-conflict and attribution costs. The 2026 METR follow-up’s time-accounting difficulty is a reminder that more capable agents do not automatically make productivity easier to measure.

For an adoption decision, compare the value of additional accepted work with the full cost: net ROI = value of additional accepted work − tool and training costs − review and rework − security and compliance costs − maintenance cost. Tool access is only one input. The relevant unit is a particular human-plus-tool workflow, applied to a particular mix of work under a particular quality standard.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Conclusion

The evidence does not support either “LLMs make experienced developers faster” or “LLMs make them slower” as a universal rule. The 2025 METR trial found a real slowdown in a narrow, demanding setting; its 2026 follow-up was consistent with improvement but could not reliably size it, while survey respondents reported perceived gains that are not objective productivity measures. For engineering leaders, the sound decision is to evaluate specific tools on representative tasks and judge quality-adjusted delivery, value and lifecycle cost together.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.