“Custom actions” is not one browser-automation API. It can mean a coordinated keyboard, pointer and wheel sequence; a reusable helper in Selenium or another framework; a Selenium IDE plugin command; a browser-extension shortcut; a WebDriver protocol extension; or, in Playwright, a custom selector engine. Choose the layer that matches what must be extended, then implement synchronization, permissions and lifecycle handling at that layer.
What “custom actions” means in browser automation
Start by identifying the object you are extending:
| Layer | Use it for | Main responsibility |
|---|---|---|
| Input action sequence | Keyboard, mouse, pen, touch or wheel gestures | Build and synchronize device events |
| Framework helper | A reusable domain operation such as “complete checkout” | Wrap locators, waits and assertions in maintainable code |
| Selenium IDE plugin | New IDE commands, locators or test-run setup | Implement the plugin lifecycle and playback contract |
| WebDriver protocol extension | A remote command unavailable in the standard protocol | Define a vendor-namespaced endpoint and remote steps |
| Chrome extension command | Keyboard shortcuts that invoke extension behavior | Declare manifest commands, events and permissions |
| Playwright selector engine | Custom element lookup semantics | Implement query and queryAll before page creation |
These layers are not interchangeable. A selector engine finds elements; it does not create a new mouse gesture. A Selenium Actions sequence runs through an automation session; it does not add a command to Selenium IDE. A Chrome extension command is an extension feature, not a portable WebDriver endpoint.
Build custom gestures with Selenium Actions
Selenium’s Actions API models virtual input devices: keyboard, pointer (mouse, pen or touch) and wheel. You chain operations and perform them as a sequence. Convenience methods cover common interactions, so use low-level actions only when the interaction genuinely needs them.
Python example: drag, key chord and wheel
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.common.action_chains import ActionChains
from selenium.webdriver.common.keys import Keys
driver = webdriver.Chrome()
try:
driver.get("https://example.test/editor")
canvas = driver.find_element(By.CSS_SELECTOR, "#canvas")
target = driver.find_element(By.CSS_SELECTOR, ".drop-zone")
actions = ActionChains(driver)
actions.move_to_element(canvas)
actions.click_and_hold()
actions.move_by_offset(120, 40)
actions.move_to_element(target)
actions.release()
actions.key_down(Keys.CONTROL).send_keys("s").key_up(Keys.CONTROL)
actions.scroll_by_amount(0, 600)
actions.perform()
finally:
driver.quit()
Replace the example URL and selectors with your application’s values. Keep the chain readable: locate elements first, then construct the gesture. If a drag operation is flaky, verify that the source and destination are visible, not covered by an overlay, and stable after any animation.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
#1 Best Overall
Synchronize multiple devices explicitly
When a sequence combines devices—for example, holding a pointer button while sending keyboard input—your test owns the synchronization. Do not assume the browser will infer your intended order. Break a complex chain into named phases, wait for the UI state between phases, and release every key or button even when an assertion fails. A leaked modifier can corrupt all subsequent tests in a reused session.
Prefer semantic helpers for repeated business actions
Wrap the low-level sequence in a helper such as resize_panel() or select_date(). The helper should own locators, explicit waits and cleanup; tests should express intent. Keep device-specific details inside the helper so changing from a mouse gesture to keyboard navigation does not require editing every test.
Add reusable commands to Selenium IDE
If the requirement is a new command visible in Selenium IDE, use its plugin mechanism rather than an Actions chain. IDE plugins can add commands and locators, run setup or teardown around test runs, and influence recording. During playback, the IDE invokes the plugin when execution reaches your custom command.
Design the command contract
- Choose a stable command name and document its target and value fields.
- Validate arguments before changing browser state; return an actionable error when a locator or option is invalid.
- Make setup and teardown idempotent so rerunning a test does not accumulate listeners or temporary data.
- Provide recording behavior only when the command can be represented reliably from user gestures.
Selenium IDE plugin pages have changed over time. Check the current Selenium IDE release documentation and package format before publishing a plugin; older examples may use APIs that no longer exist.
Free tools Windows power users keep installed
One-click scans. No signup required.
Use a WebDriver protocol extension for remote capabilities
The WebDriver 2 working draft permits additional commands that integrate with the protocol, including vendor-specific browser functions or automation of new web-platform features. This is a protocol extension, not a local helper. It is appropriate when a remote end and client need a dedicated endpoint that can be called across a grid or service boundary.
Rank #2
Namespace the endpoint
Follow the draft’s URI-template guidance: vendor-specific templates should begin with path segments that uniquely identify the vendor and user agent. Define the HTTP method, parameters, return value and error mapping, then implement the command in the remote end and expose a client wrapper. Treat the May 28, 2026 document as a working draft, not a final Recommendation; verify the current specification and browser implementation before depending on it.
When not to use a protocol extension
- For one test suite, a framework helper is simpler and more portable.
- For a keyboard shortcut inside your own extension, use the extension commands API.
- For element discovery semantics in Playwright, implement a selector engine.
Declare Chrome extension commands and shortcuts
Chrome extension commands bind keyboard shortcuts to extension behavior. Declare a commands key in the manifest, register a command event handler, and request any permissions required by the APIs your handler calls.
{
"manifest_version": 3,
"name": "Automation Helper",
"version": "1.0.0",
"background": { "service_worker": "background.js" },
"commands": {
"capture-state": {
"suggested_key": { "default": "Ctrl+Shift+Y" },
"description": "Capture the current page state"
}
},
"permissions": ["activeTab"]
}
chrome.commands.onCommand.addListener(async (command) => {
if (command !== "capture-state") return;
const [tab] = await chrome.tabs.query({active: true, lastFocusedWindow: true});
if (tab?.id) {
await chrome.tabs.sendMessage(tab.id, {type: "capture-state"});
}
});
The suggested shortcut is not guaranteed to remain unchanged: users can remap extension shortcuts in Chrome’s extension-shortcuts UI. Test command behavior with the extension installed, including cases where the active tab is a restricted page or the content script is unavailable.
Extend Playwright carefully: selector engines, not a general action registry
Playwright’s documented extension point in this area is a custom selector engine. Supply query and queryAll, then register the engine before creating a page.
import { chromium } from 'playwright';
const engine = {
query(root, selector) {
return root.querySelector(`[data-test-id="${CSS.escape(selector)}"]`);
},
queryAll(root, selector) {
return [...root.querySelectorAll(`[data-test-id="${CSS.escape(selector)}"]`)];
}
};
const browser = await chromium.launch();
const context = await browser.newContext();
await context.selectors.register('testid', engine, {contentScript: true});
const page = await context.newPage();
await page.goto('https://example.test');
await page.locator('testid=submit').click();
await browser.close();
Content-script mode can isolate the engine from page JavaScript global-object tampering while retaining DOM access. Isolation is not guaranteed when combined with other custom engines, so test the exact combination you deploy.
Rank #3
Testing browser extensions with Playwright
Use Playwright’s bundled Chromium with a persistent context when loading an extension. A persistent profile supplies the extension runtime and storage that a temporary, non-persistent context lacks. Keep the launch code in a fixture so every test receives the same profile setup. Chrome and Edge removed command-line flags needed for side-loading extensions, so do not assume a system-installed browser behaves like bundled Chromium; verify current stable documentation and versions.
Choose the right approach
| Question | Best starting point |
|---|---|
| Does the feature simulate coordinated user input? | Selenium Actions or a framework’s native input API |
| Must non-developers insert it as an IDE step? | Selenium IDE plugin |
| Does a remote browser need a new HTTP command? | WebDriver protocol extension |
| Is the behavior triggered by an extension shortcut? | Chrome commands API |
| Is the problem locating elements by a domain-specific rule? | Playwright custom selector engine |
Compare portability across browser vendors, the browser or framework control you have, synchronization and lifecycle ownership, required extension permissions, and whether a persistent profile is necessary. A local Selenium sequence and a protocol command may both be called “custom actions,” but they have different deployment and failure boundaries.
Troubleshooting common failures
Pointer action lands on the wrong element
Check viewport size, scrolling, sticky headers and overlays. Wait for the target to be visible and stable, scroll it into view, and prefer element-relative movement over hard-coded coordinates.
Keys remain pressed after a failed test
Put releases in a cleanup block. Start each test with a known input state, and recreate the driver when a session may contain leaked device state.
Actions execute out of order
Split a multi-device chain, add state-based waits, and call perform() only after the intended sequence is complete. Synchronization across devices is the caller’s responsibility.
IDE says a command is unknown
Confirm the plugin is installed and loaded in the same IDE profile, check the exact command name and argument schema, and compare your implementation with the current IDE release rather than an older example.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Rank #4
Chrome shortcut does nothing
Inspect the extension’s command registration, verify the shortcut in Chrome’s extension-shortcuts page, and test a different key combination. Restricted pages and missing permissions can prevent the handler from reaching the tab.
Playwright cannot load the extension
Use bundled Chromium, a persistent context and a valid extension path. If you are launching Chrome or Edge, account for their removal of the side-loading flags relied on by older recipes.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Or skip the browser setup
For screenshot capture rather than interactive browser control, ScreenshotNeo provides a single request that returns PNG, JPEG, WebP or PDF. It accepts cookie and consent banners like a visitor, then removes more than 60 known consent platforms, newsletter popups and chat widgets before capture; each cleanup step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers report the page verdict and billing status.
cURL:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
See the complete parameter reference at ScreenshotNeo documentation. It also offers an MCP server with take_screenshot, get_page_info and capture_pdf for Claude, Cursor and other MCP clients. Free accounts include 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →FAQ
Are custom actions portable between Selenium, Playwright and WebDriver?
No. Each framework exposes different extension points. Portability requires wrapping behavior behind your own interface and implementing that interface separately.
Should I use coordinates for a custom gesture?
Only when the interface itself is coordinate-based, such as a canvas. Element-relative actions and semantic locators generally survive layout changes better.
Best Value
Does a Playwright selector engine add a new browser command?
No. It changes how Playwright resolves locators; it does not extend the WebDriver protocol or create an extension shortcut.
Can users change a Chrome extension command’s shortcut?
Yes. Chrome exposes extension shortcuts for user remapping, so treat the manifest suggestion as a default rather than a permanent binding.
Frequently Asked Questions
Are custom actions portable between Selenium, Playwright and WebDriver?
No. Each framework exposes different extension points. Portability requires wrapping behavior behind your own interface and implementing that interface separately.
Should I use coordinates for a custom gesture?
Only when the interface itself is coordinate-based, such as a canvas. Element-relative actions and semantic locators generally survive layout changes better.
Does a Playwright selector engine add a new browser command?
No. It changes how Playwright resolves locators; it does not extend the WebDriver protocol or create an extension shortcut.
Can users change a Chrome extension command’s shortcut?
Yes. Chrome exposes extension shortcuts for user remapping, so treat the manifest suggestion as a default rather than a permanent binding.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

