A regular expression, or regex, is a small pattern language for finding and
checking text: does this input look like a date, where is every URL in this page,
which lines of the log are errors. The reference below is grouped by what you are
trying to do, and the filter box searches all of it at once. Type lookbehind, or
ES2025 to see what the newest edition added.
Regex comes in flavours that mostly agree and occasionally do not. This sheet
documents the JavaScript flavour, as defined by ECMAScript 2025, and every
pattern was run on Node.js 24 against text that should match and text that should
not. How Python and PCRE differ covers Python's
re module and PCRE2, the engine behind PHP's preg_ functions and grep -P.
For the string methods the patterns plug into, see the
JavaScript cheat sheet. For the pattern
attribute on form fields, see the HTML cheat sheet,
and for re.sub and friends, the Python cheat sheet.
Searches the task, the command and the third column. Press / from anywhere on the page.
115 commands
Using regex in JavaScript
A regex literal goes between two slashes, with any flags after the closing one. Build one with new RegExp when part of the pattern is only known while the code runs. s stands for your string and re for a regex.
| Task | Code | Notes |
|---|---|---|
| Write a regex literal | /cat/i | Flags go after the closing slash. This one ignores case |
| Build a regex from a string | new RegExp("\\d+", "g") | Every backslash is doubled inside a string: "\\d" gives \d |
| Put user input in a pattern safely | new RegExp(RegExp.escape(term), "i") | ES2025. Without escape, a dot or bracket in the input becomes regex syntax |
| Does it match | /^\d+$/.test(s) | true or false |
| First match and its groups | s.match(/(\d+)-(\d+)/) | An array: [0] is the whole match, [1] and [2] the groups. null when nothing matches |
| Every match | s.match(/\d+/g) | An array of strings, without the groups. null, not an empty array, when nothing matches |
| Every match with its groups | [...s.matchAll(/(\w+)=(\w+)/g)] | ES2020. Each item is a match array. Throws a TypeError without the g flag |
| Position of the first match | s.search(/\d/) | -1 when nothing matches |
| Replace the first match | s.replace(/cat/, "dog") | |
| Replace every match | s.replace(/cat/g, "dog") | replaceAll does the same, and throws a TypeError if the g is missing |
| Reuse the match in the replacement | s.replace(/(\w+) (\w+)/, "$2 $1") | $1 is group 1, $& the whole match, $<name> a named group, $$ a literal dollar sign |
| Work out each replacement in code | s.replace(/\d+/g, (m) => m * 2) | The function gets the match, then each group, then the position |
| Split on a pattern | s.split(/\s*,\s*/) | A capture group in the pattern puts the separators in the result too |
| Step through matches one at a time | while ((m = re.exec(s)) !== null) {} | With g or y, each call starts at re.lastIndex. null at the end, and lastIndex goes back to 0 |
| Where each group matched | /(?<year>\d{4})/d.exec(s).indices | ES2022, the d flag. [start, end] pairs, with indices.groups for named groups |
| The pattern and flags as text | re.source re.flags | /a\.b/gi gives a\.b and gi |
| An invalid pattern | new RegExp("(") | Throws a SyntaxError. A broken literal stops the whole script from loading |
Characters and escaping
Most characters match themselves. These ones mean something else: . * + ? ^ $ ( ) [ ] { } | \ and, inside a literal, /. Put a backslash in front of one to match it as an ordinary character.
| Task | Pattern | Notes |
|---|---|---|
| Some exact text | /cat/ | Anywhere in the string, so concatenate matches too |
| Any character except a line break | /c.t/ | cat, cot, c7t. With the s flag it matches a line break as well |
| A literal dot | /3\.14/ | Without the backslash, 3x14 would match too |
| Other special characters | /\(\d+\)/ | Matches (42) |
| A forward slash | /\/api\// | Escaped in a literal. new RegExp("/api/") needs no backslash |
| Tab, newline, carriage return | /\t/ /\n/ /\r/ | |
| A line break of either kind | /\r?\n/ | Windows ends lines with \r\n, macOS and Linux with \n |
| A character by its hex code | /\x41/ /é/ | A and é |
| A character past U+FFFF, such as an emoji | /\u{1F431}/u | Needs the u or v flag. Without it the pattern matches the text u{1F431} instead |
| Escape text to use as a pattern | RegExp.escape("$9.99") | ES2025. Gives \$9\.99. A leading letter or digit comes out as a \x code |
Character classes
Square brackets match one character from a set. Inside them most special characters are ordinary. The exceptions are ] and \, ^ when it comes first, and - between two characters.
| Task | Pattern | Notes |
|---|---|---|
| One of these characters | /gr[ae]y/ | gray or grey. A class matches one character, never a word |
| Any character not in the set | /[^0-9]/ | The ^ only negates when it comes first |
| A range | /[a-z]/ | a to z only. Not A, and not é |
| Several ranges | /^[a-zA-Z0-9_]+$/ | |
| A dash inside a class | /[\w.\-]/ | Escape it. The v flag, used by the HTML pattern attribute, rejects an unescaped one |
| Any digit | /\d/ | 0 to 9 only, even with the u flag. The same as [0-9] |
| Anything but a digit | /\D/ | |
| A word character | /\w/ | [A-Za-z0-9_]. No accented letters, even with the u flag |
| Anything but a word character | /\W/ | |
| Whitespace | /\s/ | Space, tab, line breaks, and Unicode spaces such as the non-breaking space |
| Anything but whitespace | /\S/ | |
| Any character, line breaks included | /[\s\S]/ | Or a dot with the s flag |
| A letter in any language | /\p{L}/u | ES2018. Needs the u or v flag. Matches é, ß, ж and 字 |
| An uppercase letter in any language | /\p{Lu}/u | |
| A letter from one script | /\p{Script=Greek}/u | Also Latin, Cyrillic, Arabic, Han and the rest |
| An emoji, skin tones and all | /^\p{RGI_Emoji}$/v | ES2024, v flag only. \p{Emoji} with u matches single code points, and the digits 0 to 9 |
| A letter, minus some | /[\p{L}--[a-z]]/v | ES2024 set subtraction: any letter except a to z |
| In both sets | /[\p{Script=Greek}&&\p{L}]/v | ES2024 intersection: Greek letters, not Greek punctuation |
Shorthands keep their meaning inside a class, so [\d,] matches a digit or a comma. Their capitals are the opposites: \D, \W and \S.
Anchors and word boundaries
Anchors match a position, not a character. They are how you say the whole string must match, rather than some part of it.
| Task | Pattern | Notes |
|---|---|---|
| At the start | /^Hello/ | |
| At the end | /world$/ | The very end. A trailing newline does not count, unlike in Python and PCRE |
| The whole string and nothing else | /^\d{5}$/ | Without ^ and $, 123456 would pass, because it contains five digits |
| Start and end of every line | /^TODO/m | The m flag makes ^ and $ match at each line break |
| A whole word | /\bcat\b/ | cat in the cat sat. Not in concatenate or cats |
| Inside a word only | /\Bcat\B/ | The cat in concatenate |
| A whole word in any language | /(?<![\p{L}\p{N}_])café(?![\p{L}\p{N}_])/u | \b only knows A to Z, 0 to 9 and _, so /\bcafé\b/ never matches café |
JavaScript has no \A or \z. Use ^ and $ without the m flag, which only ever match at the very start and end.
Quantifiers: greedy and lazy
A quantifier repeats whatever comes just before it: one character, a class or a group. Quantifiers are greedy. They take as much as they can, and give back only what the rest of the pattern needs. A ? after one makes it lazy, taking as little as it can.
| Task | Pattern | Notes |
|---|---|---|
| Zero or more | /^ab*c$/ | ac, abc, abbbc |
| One or more | /^ab+c$/ | abc, abbbc. Not ac |
| Optional: zero or one | /colou?r/ | color or colour |
| Exactly n times | /^\d{4}$/ | |
| n or more times | /^\d{2,}$/ | |
| Between n and m times | /^\d{2,4}$/ | Both ends included |
| Repeat a group | /^(ha)+$/ | ha, haha, hahaha. Without the brackets, + repeats only the a |
| Greedy: as much as possible | /<.+>/ | On <b>hi</b> it takes the whole string, from the first < to the last > |
| Lazy: as little as possible | /<.+?>/ | On <b>hi</b> it takes just <b> |
| Every lazy form | *? +? ?? {2,5}? | |
| Clearer than lazy | /<[^>]+>/ | A negated class cannot run past the >, so there is nothing to give back |
| Never give back: possessive and atomic | a++ (?>a+) | Not in JavaScript. PCRE has both, and Python since 3.11 |
Lazy does not mean shortest. /a.*?b/ on xaab matches aab, not ab, because a match starts at the first place it can.
Groups, alternation and backreferences
| Task | Pattern | Notes |
|---|---|---|
| This or that | /cat|dog/ | | has the lowest priority of anything in a pattern |
| Alternatives as a whole string | /^(?:cat|dog)$/ | Without the group, ^cat|dog$ means ^cat or dog$, so catfish and hotdog pass |
| Capture group | /(\d{4})-(\d{2})/ | Numbered from 1, in the order their opening brackets appear |
| Group without capturing | /^(?:\d{3}-)?\d{4}$/ | Groups for a quantifier or |, without taking a number |
| Named group | /(?<year>\d{4})-(?<month>\d{2})/ | ES2018. Read it as match.groups.year |
| One name in two alternatives | /(?<y>\d{4})-\d\d|\d\d\/(?<y>\d{4})/ | ES2025. Allowed because only one side can take part in a match |
| The same text again | /(\w)\1/ | A backreference. \1 matches whatever group 1 matched: the ll in hello |
| The same text again, by name | /(?<q>["']).*?\k<q>/ | Text in quotes, where the closing quote must match the opening one |
| Groups in the replacement | s.replace(/(\d+)-(\d+)-(\d+)/, "$3/$2/$1") | 2026-09-27 becomes 27/09/2026 |
| A named group in the replacement | s.replace(/(?<y>\d+)-(?<m>\d+)/, "$<m>/$<y>") | |
| The whole match in the replacement | s.replace(/\d+/g, "[$&]") | a1b2 becomes a[1]b[2] |
| A dollar sign in the replacement | s.replace(/price/, "$$5") | $$ gives one $. Matters when the replacement text comes from users |
Nested groups count too, by opening bracket: in /((a)(b))/, group 1 is ab, group 2 is a and group 3 is b.
Lookahead and lookbehind
A lookaround checks what comes before or after the current position without making it part of the match. Like ^ and \b, it matches a position rather than any characters.
| Task | Pattern | Notes |
|---|---|---|
| Followed by | /\d+(?=px)/ | The 12 in 12px, without the px |
| Not followed by | /q(?!u)/i | A q with no u after it: Iraq, qat. Not queen |
| Preceded by | /(?<=£)\d+/ | ES2018. The 30 in £30, and nothing in $30 |
| Not preceded by | /(?<!-)\b\d+/g | Positive whole numbers only: 12 and 7 in 12 -5 7 |
| Must contain several things, in any order | /^(?=.*\d)(?=.*[a-z]).{8,}$/ | Each lookahead scans from the start: a digit, a lowercase letter, 8 characters or more |
| Must not contain | /^(?!.*password).*$/i | A line that does not include password anywhere |
| Add thousands separators | s.replace(/\B(?=(\d{3})+(?!\d))/g, ",") | 1234567 becomes 1,234,567. For display, toLocaleString does it for every locale |
| Split before each capital | s.split(/(?=[A-Z])/) | camelCaseName becomes camel, Case, Name |
A JavaScript lookbehind can hold any pattern, + and * included. Python only allows a fixed length, and PCRE2 a bounded one. Safari has supported lookbehind since 16.4, released in March 2023.
Flags
| Task | Example | Notes |
|---|---|---|
| g: every match, not just the first | /cat/g | Needed by matchAll and replaceAll. Also makes test and exec remember where they stopped |
| i: ignore case | /hello/i | |
| m: ^ and $ at every line | /^\d+$/m | Multiline |
| s: dot matches line breaks too | /<p>.*<\/p>/s | ES2018. Also called dotAll |
| u: Unicode mode | /^.$/u | Whole characters, so . matches all of an emoji. Allows \p{...}, and turns typos like \q into errors |
| v: Unicode sets | /[\p{L}--[a-z]]/v | ES2024. A stricter u that adds set operations and \p{RGI_Emoji}. Use u or v, not both |
| y: match only at lastIndex | /\d+/y | Sticky. For tokenisers that walk a string from left to right |
| d: start and end of each group | /(?<n>\d+)/d | ES2022. Adds indices to the match |
| Several flags | /cat/gi | In any order |
| A flag for part of the pattern | /(?i:hello) World/ | ES2025. i, m and s only. (?-i:...) turns one off |
| Which flags a regex has | re.flags re.global re.ignoreCase | A string like gi, then true or false for each |
Python and PCRE also have x, verbose mode, for spreading a pattern over lines with comments. JavaScript has no equivalent.
Common patterns
Starting points, each run against text that should match and text that should not. They check the shape of the text, not whether it is real: no pattern can tell you an email address exists or a date happened. Read the caveat before you copy one.
| Task | Pattern | Caveats |
|---|---|---|
| Email address, rough check | /^[^\s@]+@[^\s@]+\.[^\s@]+$/ | Something, @, something, a dot, something. Lets a@b.c through. The only real check is sending an email |
| URL starting with http or https | /^https?:\/\/[^\s/$.?#][^\s]*$/i | A rough check. URL.canParse(s) is stricter, and new URL(s) also splits it into parts |
| URLs inside a piece of text | /https?:\/\/[^\s<>"]+/g | Also grabs a full stop or bracket that ends the sentence. Trim those afterwards |
| ISO date, shape only | /^\d{4}-\d{2}-\d{2}$/ | Also accepts 2026-99-99 |
| ISO date, real months and days | /^\d{4}-(?:0[1-9]|1[0-2])-(?:0[1-9]|[12]\d|3[01])$/ | Still accepts 2026-02-31. Check the day in code: see the dates section below |
| Day/month/year date | /^(?:0[1-9]|[12]\d|3[01])\/(?:0[1-9]|1[0-2])\/\d{4}$/ | dd/mm/yyyy. 03/04/2026 is 3 April in the UK and 4 March in the US, and no pattern can tell which was meant |
| 24-hour time | /^(?:[01]\d|2[0-3]):[0-5]\d$/ | 00:00 to 23:59, with the leading zero |
| Whole number | /^-?\d+$/ | Also allows leading zeros, like 007 |
| Decimal number | /^-?\d+(?:\.\d+)?$/ | Not .5, 1e3 or 1,000. Number(s) accepts more forms |
| Hex colour | /^#(?:[\da-f]{3,4}|[\da-f]{6}|[\da-f]{8})$/i | #fff, #ffff, #ffffff and #ffffff80, with or without alpha |
| URL slug | /^[a-z0-9]+(?:-[a-z0-9]+)*$/ | my-first-post. No capitals, and no dash at either end or twice in a row |
| Username | /^[a-zA-Z0-9_]{3,16}$/ | Letters, digits and underscores, 3 to 16 long. ASCII only, so José fails |
| IPv4 address | /^(?:(?:25[0-5]|2[0-4]\d|1\d\d|[1-9]?\d)\.){3}(?:25[0-5]|2[0-4]\d|1\d\d|[1-9]?\d)$/ | 0 to 255 in each part, and no leading zeros |
| UUID | /^[\da-f]{8}-[\da-f]{4}-[\da-f]{4}-[\da-f]{4}-[\da-f]{12}$/i | Any version. crypto.randomUUID() makes version 4 ones |
| UK postcode | /^[A-Z]{1,2}\d[A-Z\d]? ?\d[A-Z]{2}$/i | Shape only: SW1A 1AA and M1 1AE pass, and so do codes that were never issued |
| Password with a mix of characters | /^(?=.*[a-z])(?=.*[A-Z])(?=.*\d).{12,}$/ | Length does more than character rules. Check against leaked-password lists too |
| Blank or only whitespace | /^\s*$/ | |
| Collapse runs of whitespace | s.replace(/\s+/g, " ").trim() | |
| A word typed twice | /\b(\w+)\s+\1\b/i | Finds the the. The i flag also catches The the |
| Strip HTML tags | s.replace(/<[^>]*>/g, "") | For your own simple markup. Not safe for sanitising user HTML: use DOMPurify or the browser's DOMParser |
Named groups and matchAll, start to finish
Most of the tables in one place: named groups, the g and m flags, matchAll, a
lookbehind and a replacer function, used to pull a log apart.
const log = `2026-09-27 14:02:11 ERROR [payments] Card declined for order 1042
2026-09-27 14:02:15 INFO [search] 12 results for "cat beds"
2026-09-27 14:03:40 WARN [payments] Retry 2 of 3 for order 1042
2026-09-27 14:05:02 ERROR [auth] Token expired for user 77`;
const line =
/^(?<date>\d{4}-\d{2}-\d{2}) (?<time>\d{2}:\d{2}:\d{2}) (?<level>ERROR|WARN|INFO) \[(?<service>[\w-]+)\] (?<message>.*)$/gm;
const entries = [...log.matchAll(line)].map((m) => ({ ...m.groups }));
console.log(entries.length); // 4
console.log(entries[1].message); // 12 results for "cat beds"
const errors = entries
.filter((e) => e.level === "ERROR")
.map((e) => `${e.time} ${e.service}`);
console.log(errors); // [ '14:02:11 payments', '14:05:02 auth' ]
const orders = new Set(log.match(/(?<=order )\d+/g));
console.log([...orders]); // [ '1042' ]
const redacted = log.replace(/(?<=user )\d+/g, (id) => "*".repeat(id.length));
console.log(redacted.split("\n").at(-1)); // 2026-09-27 14:05:02 ERROR [auth] Token expired for user **The m flag is what lets ^ and $ match at each line of the log rather than
only at its very start and end, and matchAll needs the g. .* in the message
group cannot run on to the next line, because a dot does not match a line break.
The lookbehind in (?<=order )\d+ checks for the word without including it, so the
match is just the number. Copying m.groups into a plain object with { ...m.groups }
is optional, but it prints more tidily: the original has no prototype.
The HTML pattern attribute, start to finish
A form field can check its value against a regex with no JavaScript at all.
<label for="ref">Order reference</label>
<input id="ref" name="ref" required
pattern="[A-Z]{3}-\d{4}"
title="Three capital letters, a dash and four digits, like CAT-2026">The browser wraps the attribute in ^(?: and )$ and compiles it with the v
flag, so this is what it runs:
const pattern = String.raw`[A-Z]{3}-\d{4}`;
const re = new RegExp(`^(?:${pattern})$`, "v");
console.log(re.test("CAT-2026")); // true
console.log(re.test("CAT-2026x")); // false: the whole value has to match
console.log(re.test("cat-2026")); // false: case-sensitive, and there is no i flag
try {
new RegExp(String.raw`^(?:[\w-]+)$`, "v");
} catch (err) {
console.log(err.name); // SyntaxError, so a browser ignores pattern="[\w-]+"
}Three things follow. You never need ^ and $ in a pattern, since the whole
value always has to match. There is no way to add a flag, so allow both cases with
[A-Za-z]. And because of the v flag, a pattern with an unescaped - in a class
is invalid, and the browser skips the check without a word. The
HTML cheat sheet has the rest of the form
attributes, and the server still has to check the value again.
Email, URL and date patterns: what they miss
The common patterns table checks the shape of text. For email, URLs and dates, this is how far a pattern can go, and what to use once it runs out.
// The pattern browsers use for <input type="email">, from the HTML standard
const htmlEmail =
/^[a-zA-Z0-9.!#$%&'*+\/=?^_`{|}~-]+@[a-zA-Z0-9](?:[a-zA-Z0-9-]{0,61}[a-zA-Z0-9])?(?:\.[a-zA-Z0-9](?:[a-zA-Z0-9-]{0,61}[a-zA-Z0-9])?)*$/;
console.log(htmlEmail.test("ada@example.com")); // true
console.log(htmlEmail.test("ada@localhost")); // true: no dot needed after the @
console.log(htmlEmail.test("josé@example.com")); // false: é is not on the list
console.log(htmlEmail.test('"ada lovelace"@example.com')); // false, though email allows it
// URLs: parse them rather than pattern-match them
function isWebUrl(text) {
if (!URL.canParse(text)) return false;
const { protocol } = new URL(text);
return protocol === "https:" || protocol === "http:";
}
console.log(isWebUrl("https://example.com/a?b=1")); // true
console.log(isWebUrl("javascript:alert(1)")); // false: a valid URL, but not a web one
console.log(isWebUrl("example.com")); // false: no scheme
// Dates: the regex checks the shape, a round trip checks the day exists
function isRealIsoDate(text) {
const m = /^(?<y>\d{4})-(?<mo>\d{2})-(?<d>\d{2})$/.exec(text);
if (!m) return false;
const [y, mo, d] = [m.groups.y, m.groups.mo, m.groups.d].map(Number);
const date = new Date(Date.UTC(y, mo - 1, d));
return date.getUTCFullYear() === y && date.getUTCMonth() === mo - 1 && date.getUTCDate() === d;
}
console.log(isRealIsoDate("2026-09-27")); // true
console.log(isRealIsoDate("2026-02-31")); // false: Date rolls it over to 3 March
console.log(isRealIsoDate("2028-02-29")); // true: 2028 is a leap yearEven the standard's own email pattern is a deliberate compromise: it rejects
addresses the email standard allows, and accepts ones no mail server will deliver
to. Use it, or the rough check from the table, to catch typos, then send a
confirmation email. For URLs, URL.canParse does the parsing and the protocol check
rules out javascript: links, which a loose pattern would let through. For dates,
building a Date from the parts and reading them back is what catches 31 February,
because Date quietly rolls an impossible day over into the next month.
How Python and PCRE differ
The core is shared: classes, anchors, quantifiers, groups, | and lookahead read
the same everywhere. These are the differences that break a pattern when you copy
it from one language to another. Python is the re module in Python 3.14. PCRE is
PCRE2, which PHP's preg_ functions and grep -P use.
| Feature | JavaScript (ES2025) | Python 3.14 re | PCRE2 |
|---|---|---|---|
| Named group | (?<name>...) | (?P<name>...) only | Both |
| Backreference by name | \k<name> | (?P=name) | Both |
| Group in the replacement | $1, $<name> | \1, \g<name> | Set by the language: $1 in PHP |
| Replace every match | Needs the g flag | re.sub does by default, count=1 for one | Set by the language |
| Match the whole string | ^...$ | re.fullmatch(), or \A...\z | \A...\z |
$ before a final newline | Does not match | Matches | Matches |
\d and \w | ASCII only | Any Unicode digit or letter. re.ASCII limits them | ASCII, unless Unicode (UCP) mode is on |
\p{L} and other Unicode properties | With the u or v flag | Not supported | Supported |
| Lookbehind | Any pattern | Fixed length only | Bounded length, like a{1,3} but not a+ |
Possessive a++ and atomic (?>...) | Not supported | Since 3.11 | Supported |
| Inline flags | Scoped only: (?i:...) | (?i) at the very start, or scoped | (?i) anywhere, or scoped |
| Comments and whitespace in the pattern | Not supported | re.X | (?x) |
| One group name used twice | In separate alternatives | Error | Error, unless (?J) |
| Recursion, for nested brackets | Not supported | Not supported | (?R) and (?1) |
| Escape text for a pattern | RegExp.escape() | re.escape() | preg_quote() in PHP |
Two Python habits trip people up the other way round. re.match only looks at the
start of the string, so re.match(r"\d+", "ab12") is None: use re.search to
look anywhere. And re.findall returns the groups rather than the whole match when
the pattern has any, so re.findall(r"(\w+)@\w+", s) gives just the names before
the @.
Gotchas
The regex mistakes almost everyone makes, most of them more than once.
| Looks right | What actually happens | Do this instead |
|---|---|---|
/a/g saved in a variable, then re.test(s) in a loop | Alternates true and false on a string with one match: g makes test carry on from lastIndex | Drop the g for test, or set re.lastIndex = 0 first |
new RegExp("\d+") | "\d" in a string is just d, so this matches ddd | new RegExp("\\d+"), or String.raw`\d+` |
new RegExp(userInput) | A . or ( in the input becomes syntax, or throws a SyntaxError | new RegExp(RegExp.escape(userInput)) |
/".*"/ on "a" and "b" | Matches the whole thing, from the first quote to the last | /"[^"]*"/, or the lazy /".*?"/ |
/[A-z]/ | Also matches [, \, ], ^, _ and a backtick, which sit between Z and a | /[A-Za-z]/ |
/^cat|dog$/ | Matches catfish and hotdog: | splits the whole pattern | /^(?:cat|dog)$/ |
/3.14/ | Also matches 3x14: a dot is any character | /3\.14/ |
. to match anything | Stops at a line break | [\s\S], or the s flag |
s.match(/(\w)=(\d)/g) for the groups | With g, match returns whole matches only | [...s.matchAll(/(\w)=(\d)/g)] |
s.replace(re, userText) | A $& or $1 typed by the user is expanded | s.replace(re, () => userText) |
/^(a+)+$/ | Nested quantifiers: 30 a's and a ! took about 5 seconds on Node.js 24 | /^a+$/. Never nest a quantifier inside another over the same text |
/\bcafé\b/ | Never matches: \b treats é as a non-word character | Lookarounds with \p{L} and the u flag, as in the anchors table |
/^.{1,10}$/ to limit a length | Without u, each emoji counts as two | Add the u flag |
pattern="[\w-]+" in HTML | Invalid under the v flag browsers use, so the check is skipped | pattern="[\w\-]+" |
re.search(r"^\d+$", "42\n") in Python | Matches: Python's $ allows a final newline | re.fullmatch(r"\d+", s), or \z in Python 3.14 |
| A regex to parse HTML or JSON | Breaks on nesting, comments and quoted > characters | DOMParser or JSON.parse |
Common questions
Which regex flavour does this cheat sheet cover?
JavaScript regular expressions as defined by ECMAScript 2025, the current edition of the standard, with every pattern run on Node.js 24. Most of the syntax, such as classes, anchors, quantifiers and groups, is the same in Python, PHP, Java and C#. The section on how Python and PCRE differ lists the places where it is not, checked on Python 3.14 and PCRE2 10.47.
What is the difference between greedy and lazy quantifiers?
A greedy quantifier such as * or + takes as many characters as it can, then gives some back if the rest of the pattern needs them. A lazy one, written with a ? after it such as *? or +?, takes as few as it can and adds more only when it has to. On the text <b>hi</b>, the greedy /<.+>/ matches the whole string and the lazy /<.+?>/ matches just <b>. A negated class like /<[^>]+>/ is often clearer than either.
What does ?: mean in a regex?
(?:...) is a non-capturing group. It groups part of a pattern so a quantifier or | applies to all of it, as in (?:ab)+ or ^(?:cat|dog)$, without saving what it matched or taking a group number. Use it whenever you need brackets but not the captured text. The other ? forms are lookarounds, (?=...), (?!...), (?<=...) and (?<!...), and named groups, (?<name>...).
How do I match a literal dot, bracket or other special character?
Put a backslash in front of it: \. matches a dot, \( an opening bracket and \\ a backslash. The characters that need it are . * + ? ^ $ ( ) [ ] { } | and the backslash, plus / inside a regex literal. To use text you do not control, such as a search box, as part of a pattern, pass it through RegExp.escape first. That is ES2025, available in Node.js 24 and browsers released since 2025.
What is the best regex for validating an email address?
There is no pattern that accepts every valid address and rejects every invalid one, and no pattern can tell you the address exists. A rough check such as /^[^\s@]+@[^\s@]+\.[^\s@]+$/ catches typos, and input type="email" gives you the browser's own check for free. The only real test is sending a confirmation email and waiting for someone to click the link.
What do the g, i and m flags do?
g finds every match instead of stopping at the first, and is needed by matchAll and replaceAll. i ignores case, so /hello/i matches HELLO. m makes ^ and $ match at the start and end of every line rather than only the whole string. They combine in any order, as in /^todo/gim. The other flags are s, so a dot matches line breaks, u and v for Unicode, y for sticky matching and d for match positions.
Why does my regex work in Python but not in JavaScript?
Usually one of five things. Python writes named groups as (?P<name>...), which JavaScript rejects; JavaScript uses (?<name>...). Python's \d and \w match any Unicode digit or letter, JavaScript's only ASCII. Python's $ also matches before a final newline. Python has possessive quantifiers, atomic groups and the verbose x flag, which JavaScript does not. And Python uses re.sub for every match by default, where JavaScript needs the g flag.
Why does my regex freeze the page?
Catastrophic backtracking. A pattern with a quantifier inside a quantifier, such as /^(a+)+$/, can try an exponential number of ways to split the input before it gives up on a string that nearly matches. On Node.js 24, thirty a's followed by an exclamation mark took about five seconds. Remove the nesting (/^a+$/ matches the same strings), make the alternatives inside a repeated group unable to match the same text, and limit the length of any input you run a pattern over.
Can I test these patterns without installing anything?
Yes. Open your browser's console with F12 and type a pattern with test, such as /^\d{4}$/.test("2026"), which prints true. Try a few strings that should match and a few that should not, especially the edge cases: an empty string, a trailing space, a line break. Keep those strings as unit tests beside the pattern, so the next change cannot quietly break it.
