Skip to main content
ZEKILO Dev

Character & Byte Counter

Paste text to see its character count and its size in bytes for each encoding.

Stays in your browserWHATWG EncodingUAX #29As of 2026.10.03

Input

Press Esc, then Tab to leave the editor.

Counts

Character and byte counts
Characters0Grapheme clusters: an emoji sequence counts as one
Characters without whitespace0
Code points0
UTF-16 length0The length in JavaScript and Java
UTF-8 bytes0
EUC-KR (CP949) bytes0
Shift_JIS (CP932) bytes0
Lines0
Words0

How to use Character & Byte Counter

  1. 1Paste text into the input panel, or press Sample. A text file can be opened with Open file.
  2. 2Read the character count and the UTF-8, EUC-KR, and Shift_JIS byte counts in the table. Enter a character limit to see how many characters are left or over.
  3. 3Use the copy button on a row to copy that value.

Character & Byte Counter options

Line break bytes
Counts each line break as LF (1 byte, default) or CRLF (2 bytes). It affects the three byte counts only. Choose CRLF when the text will be saved on Windows or sent in a fixed-length record that uses CRLF.
Character limit
Enter a number to see how many characters are left or how many you are over, based on the character count. 0 means no limit. For example, 280 for a post on X or 1,000 for an application form.

What counts as a character

The length of a text depends on the unit you count. The character count in this tool uses grapheme clusters as defined in Unicode UAX #29, which are the units that appear as a single character on screen. 👍🏻 is 2 code points and 4 UTF-16 code units, but it counts as 1 character. The length of a string in JavaScript or Java is the number of UTF-16 code units, which is the UTF-16 length row.

Characters without whitespace leaves out spaces, tabs, line breaks, and the full-width space (U+3000). The line count is the number of line breaks plus one, and empty input has 0 lines.

Bytes in EUC-KR and Shift_JIS

EUC-KR follows CP949 (MS949) and Shift_JIS follows CP932 (Windows-31J). Other systems may differ. This is how the WHATWG Encoding Standard used by browsers defines the two names, and MS949 and Windows-31J on Windows and in Java use the same tables.

EUC-KR in the narrow sense (KS X 1001) has only 2,350 Hangul syllables, so a syllable such as 똠 is missing. CP949 adds the other 8,822 in an extension area, so all 11,172 modern syllables take 2 bytes. A system that does not support the extension may still reject those characters, so check the rules of the receiving side for fixed-length records and older databases.

In Shift_JIS a half-width katakana character such as カ takes 1 byte, and full-width kana and kanji take 2. Characters that are not in JIS X 0208, such as ① and 髙, are in the CP932 extension.

Character & Byte Counter examples

  • Hangul and Latin letters — 12 characters

    Shift_JIS has no Hangul, so 5 characters cannot be represented.

    Input

    안녕하세요 ZEKILO

    Output

    UTF-16 12 · UTF-8 22 · EUC-KR 17 · Shift_JIS —
  • Hiragana and Latin letters — 12 characters

    KS X 1001, the character set behind EUC-KR, includes kana.

    Input

    こんにちは ZEKILO

    Output

    UTF-16 12 · UTF-8 22 · EUC-KR 17 · Shift_JIS 17
  • An emoji sequence and Hangul — 2 characters, 3 code points

    The thumbs-up emoji and its skin tone modifier are one character on screen.

    Input

    👍🏻가

    Output

    UTF-16 5 · UTF-8 11 · EUC-KR — · Shift_JIS —
  • Hangul outside KS X 1001 — 4 characters

    The first syllable is not in KS X 1001, but it is in the CP949 extension and takes 2 bytes.

    Input

    똠방각하

    Output

    UTF-16 4 · UTF-8 12 · EUC-KR 8 · Shift_JIS —
  • Characters outside JIS X 0208 — 2 characters

    Both characters are in the CP932 extension. EUC-KR has no code for the second one.

    Input

    ①髙

    Output

    UTF-16 2 · UTF-8 6 · EUC-KR — · Shift_JIS 4

Examples use the reference cases this tool is tested against.

Character & Byte Counter: frequently asked questions

Is my input sent to a server?

No. Your input and the result are processed only in this browser and are not stored. Reloading the page clears them.

Which definition of EUC-KR and Shift_JIS is used for the byte counts?

EUC-KR follows CP949 (MS949) and Shift_JIS follows CP932 (Windows-31J). Other systems may differ. Hangul, kana, and kanji take 2 bytes and ASCII letters and digits take 1. If the text has characters the encoding does not contain, the row shows a dash and the number of characters that cannot be represented.

How are emoji and combined characters counted?

The character count uses grapheme clusters, the units you see as one character. An emoji with a skin tone, a flag, a family emoji, and a letter with a combining accent each count as one. Code points and UTF-16 length are shown in separate rows, so you can compare them with the length your program reports.

Why is the UTF-16 length different from the character count?

The length of a string in JavaScript and Java is the number of UTF-16 code units. Characters outside the Basic Multilingual Plane, such as most emoji, take two code units, and one visible character can be made of several code points. The thumbs-up emoji with a skin tone is 1 character, 2 code points, and 4 UTF-16 code units.

How do I check text against a length limit?

Enter the limit in the Character limit option. The number of characters left is shown above the table, and when you go over, the excess is highlighted. For a limit in bytes, such as a database column, read the byte row for the encoding that system uses.

What does a dash in a byte row mean?

The text contains characters that the encoding cannot represent. Shift_JIS has no Hangul, for example, and neither EUC-KR nor Shift_JIS has emoji. The number of such characters is shown under the row and in the note below the table.

How it works · Standards

  • Processed in: this browser (your device). Nothing is sent or stored.
  • Standards: WHATWG Encoding · UAX #29
  • Engine: ZEKILO Dev character and byte counter (in-house code) · EUC-KR and Shift_JIS encoding tables (built from the WHATWG Encoding indexes) (BSD 3-Clause (WHATWG))
  • Characters are segmented with the browser's Intl.Segmenter. In a browser without it, characters are counted as code points and the tool says so.
  • The number of characters that cannot be represented is counted in code points.
  • Words are pieces separated by whitespace, so a sentence written without spaces counts as one word.
  • As of 2026.10.03
  • Changelog: 2026.10.03 First release
Open-source licenses →
Something wrong with the result? Report an issue(Your input is not attached)