Skip to main content
ZEKILO Dev
Text

Regex Basics: 10 Patterns and How to Avoid ReDoS

Updated 2026.10.03 · 6 min read

Ten regular expression patterns with tested examples, plus how catastrophic backtracking (ReDoS) happens with patterns like (a+)+ and how to rewrite them.

Open Regex Tester

Ten building blocks cover most everyday regular expressions: the dot, three shorthand classes, character sets, anchors, quantifiers, groups and lazy matching. The one habit to avoid is a repetition inside another repetition, as in (a+)+, because a backtracking engine can need exponentially more time on text that almost matches. You can try every example on this page in the Regex Tester, which uses JavaScript (ECMAScript) syntax.

Ten basic patterns

Each example below was run with the g flag, which finds every match instead of only the first.

# Pattern Meaning Example Matches
1 . Any one character except a line break c.t in cat cot c-t ct cat, cot, c-t
2 \d A digit, 0 to 9 \d+ in v2.10 build 345 2, 10, 345
3 \w An ASCII letter, a digit or an underscore \w+ in user_id=42; name=kim user_id, 42, name, kim
4 \s White space: space, tab or line break a\s+b in a b, a b, ab a b, a b
5 [abc] [^abc] One of the listed characters, or none of them [aeiou] in zekilo e, i, o
6 ^ $ Start and end of the text \.json$ in data.json .json
7 * + ? Zero or more, one or more, zero or one colou?r in color colour colr color, colour
8 {n} {n,m} An exact count, or a range of counts a{2,3} in a aa aaa aaaa aa, aaa, aaa
9 (…) | A group, and alternation between options cat|dog in cat, dog, cow cat, dog
10 *? +? Lazy: repeat as few times as possible <.+?> in <b>bold</b> <b>, </b>

A few details are worth knowing early:

  • Escape special characters. A dot matches almost anything, so write \. for a literal dot, as in row 6.
  • Quantifiers are greedy by default. Without the ? in row 10, <.+> matches the whole text <b>bold</b> as a single match.
  • \w and \d are ASCII-only in JavaScript. \w+ finds only caf in café. With the u flag, \p{L}+ matches letters in any script and returns café.
  • Flags change the rules. i ignores case, m makes ^ and $ work on every line (^ERROR then matches at the start of each line that begins with it), and s lets the dot match line breaks.

Groups, named groups and replacing

Parentheses also capture what they match, so you can reuse the pieces. Take this pattern and text (the label 연락처 means “contact”):

Pattern   (\d{3})-(\d{4})-(\d{4})    flag g
Text      연락처: 010-1234-5678, 02-123-4567, 010-9876-5432

Match 1   010-1234-5678   at index 5    groups: 010, 1234, 5678
Match 2   010-9876-5432   at index 33   groups: 010, 9876, 5432

The middle number is skipped because its first part has only two digits. In replace mode, $1 to $3 refer to the groups, so the replacement $1-****-$3 produces:

연락처: 010-****-5678, 02-123-4567, 010-****-5432

Numbered groups become hard to follow as a pattern grows. A named group, written (?<name>…), labels the piece instead: (?<year>\d{4})-(?<m>\d{2}) applied to 2026-10 gives year = 2026 and m = 10, and a replacement can refer to them as $<year> and $<m>.

What catastrophic backtracking is

A backtracking engine tries one way of matching, and when that fails it goes back and tries the next one. Usually there are only a few ways. With (a+)+$ there are a great many: the inner a+ and the outer + can split a run of the letter a into groups in every possible combination. If the text ends with a character that makes the match fail, the engine works through all of those combinations before it can report that there is no match.

OWASP’s description of the problem uses ^(a+)+$ as the example: for the input aaaaX there are 16 possible paths, for 16 letters followed by X there are 65,536, and the number doubles with each additional a. A timing run shows the same shape. Testing (a+)+$ against a run of the letter a followed by b in Node.js 24.14 on one desktop PC took:

Length of the run Time (median of 3 runs)
20 20 ms
22 85 ms
24 300 ms
26 1,311 ms

Your numbers will differ by machine and engine, but the growth of roughly four times for every two extra characters will not. On the same 26-letter text, a+$, which matches exactly the same text without the nested repetition, returned in less than 0.1 ms.

This is called ReDoS, regular expression denial of service. OWASP defines it as a denial-of-service attack that exploits the fact that most regular expression implementations may reach extreme situations that cause them to work very slowly. If a server applies such a pattern to text that a visitor controls, one request can keep a CPU core busy for a very long time.

OWASP lists two ingredients of a risky pattern: a group with repetition, and inside that group either another repetition or alternatives that overlap. Its examples are (a+)+$, ([a-zA-Z]+)*$, (a|aa)+$ and (a|a?)+$.

How to avoid it

  1. Do not nest quantifiers. (a+)+ matches the same text as a+, and (\d+)* the same as \d*. If a group is repeated, make sure the group cannot match the same text in more than one way.
  2. Keep alternatives from overlapping. In (a|aa)+ both options can start at the same place. Rewrite it so that each position has one possible reading; here, a+.
  3. Prefer a narrow class to .*. To match a quoted string, "[^"]*" finds "a" and "b" separately in "a" and "b", while ".*" grabs everything from the first quote to the last and has to back off character by character whenever the rest of a longer pattern fails.
  4. Limit the input length and anchor the pattern. Even a+$ is not free on very long text: without an anchor at the start, the search restarts from every position. On 100,000 letters followed by b it took about 7 seconds in the same setup, while ^a+$ took less than 1 ms. Check the length of a field before you run a pattern on it.
  5. Run with a time limit. The Regex Tester always runs your pattern in a separate worker and stops it after 1,000 ms, so (a+)+$ against 30 letters plus b ends with the message “The pattern takes too long” instead of freezing the page.
  6. Do not build patterns from user input. If users must supply their own patterns, consider an engine that guarantees linear-time matching, such as RE2. Its README explains the trade-off: it does not support backreferences or look-around assertions.

One JavaScript-specific note: advice written for PCRE or Java often suggests atomic groups (?>…) or possessive quantifiers such as a++. As of October 2026 these are not part of the ECMAScript regular expression grammar, and Node.js 24 rejects both as syntax errors, so in JavaScript the fix is to rewrite the pattern.

Testing a pattern with ZEKILO Dev

  1. Open the Regex Tester, enter the pattern and paste the text to search.
  2. Choose flags (g, i, m, s, u, y, d, plus v in browsers that support it) and a mode: find, replace or split.
  3. Read the match count and the table of matches with their positions, groups and named groups. Matches are also highlighted in the text.

Before you ship a pattern, try it on text that almost matches, for example a long valid value with one wrong character at the end. If the tester reports a timeout, the pattern needs rewriting. The pattern and text stay in your browser and are not sent to a server.

Summary

  • Ten pieces go a long way: ., \d, \w, \s, character sets, anchors, quantifiers, counts, groups with alternation, and lazy matching.
  • Capture groups and named groups let you reuse parts of a match in a replacement.
  • Nested or overlapping repetition such as (a+)+ can take exponentially long on text that almost matches.
  • Rewrite such patterns, limit input length, and run untrusted patterns with a time limit.

Tools for this guide

Sources

More guides