Character & Byte Counter
Paste text to see its character count and its size in bytes for each encoding.
Stays in your browserWHATWG EncodingUAX #29As of 2026.10.03
Input
Counts
| Characters | 0Grapheme clusters: an emoji sequence counts as one | |
|---|---|---|
| Characters without whitespace | 0 | |
| Code points | 0 | |
| UTF-16 length | 0The length in JavaScript and Java | |
| UTF-8 bytes | 0 | |
| EUC-KR (CP949) bytes | 0 | |
| Shift_JIS (CP932) bytes | 0 | |
| Lines | 0 | |
| Words | 0 |
How to use Character & Byte Counter
- 1Paste text into the input panel, or press Sample. A text file can be opened with Open file.
- 2Read the character count and the UTF-8, EUC-KR, and Shift_JIS byte counts in the table. Enter a character limit to see how many characters are left or over.
- 3Use the copy button on a row to copy that value.
Character & Byte Counter options
- Line break bytes
- Counts each line break as LF (1 byte, default) or CRLF (2 bytes). It affects the three byte counts only. Choose CRLF when the text will be saved on Windows or sent in a fixed-length record that uses CRLF.
- Character limit
- Enter a number to see how many characters are left or how many you are over, based on the character count. 0 means no limit. For example, 280 for a post on X or 1,000 for an application form.
What counts as a character
The length of a text depends on the unit you count. The character count in this tool uses grapheme clusters as defined in Unicode UAX #29, which are the units that appear as a single character on screen. 👍🏻 is 2 code points and 4 UTF-16 code units, but it counts as 1 character. The length of a string in JavaScript or Java is the number of UTF-16 code units, which is the UTF-16 length row.
Characters without whitespace leaves out spaces, tabs, line breaks, and the full-width space (U+3000). The line count is the number of line breaks plus one, and empty input has 0 lines.
Bytes in EUC-KR and Shift_JIS
EUC-KR follows CP949 (MS949) and Shift_JIS follows CP932 (Windows-31J). Other systems may differ. This is how the WHATWG Encoding Standard used by browsers defines the two names, and MS949 and Windows-31J on Windows and in Java use the same tables.
EUC-KR in the narrow sense (KS X 1001) has only 2,350 Hangul syllables, so a syllable such as 똠 is missing. CP949 adds the other 8,822 in an extension area, so all 11,172 modern syllables take 2 bytes. A system that does not support the extension may still reject those characters, so check the rules of the receiving side for fixed-length records and older databases.
In Shift_JIS a half-width katakana character such as カ takes 1 byte, and full-width kana and kanji take 2. Characters that are not in JIS X 0208, such as ① and 髙, are in the CP932 extension.
Character & Byte Counter examples
Hangul and Latin letters — 12 characters
Shift_JIS has no Hangul, so 5 characters cannot be represented.
Input
안녕하세요 ZEKILO
Output
UTF-16 12 · UTF-8 22 · EUC-KR 17 · Shift_JIS —
Hiragana and Latin letters — 12 characters
KS X 1001, the character set behind EUC-KR, includes kana.
Input
こんにちは ZEKILO
Output
UTF-16 12 · UTF-8 22 · EUC-KR 17 · Shift_JIS 17
An emoji sequence and Hangul — 2 characters, 3 code points
The thumbs-up emoji and its skin tone modifier are one character on screen.
Input
👍🏻가
Output
UTF-16 5 · UTF-8 11 · EUC-KR — · Shift_JIS —
Hangul outside KS X 1001 — 4 characters
The first syllable is not in KS X 1001, but it is in the CP949 extension and takes 2 bytes.
Input
똠방각하
Output
UTF-16 4 · UTF-8 12 · EUC-KR 8 · Shift_JIS —
Characters outside JIS X 0208 — 2 characters
Both characters are in the CP932 extension. EUC-KR has no code for the second one.
Input
①髙
Output
UTF-16 2 · UTF-8 6 · EUC-KR — · Shift_JIS 4
Examples use the reference cases this tool is tested against.
Character & Byte Counter: frequently asked questions
Is my input sent to a server?
No. Your input and the result are processed only in this browser and are not stored. Reloading the page clears them.
Which definition of EUC-KR and Shift_JIS is used for the byte counts?
EUC-KR follows CP949 (MS949) and Shift_JIS follows CP932 (Windows-31J). Other systems may differ. Hangul, kana, and kanji take 2 bytes and ASCII letters and digits take 1. If the text has characters the encoding does not contain, the row shows a dash and the number of characters that cannot be represented.
How are emoji and combined characters counted?
The character count uses grapheme clusters, the units you see as one character. An emoji with a skin tone, a flag, a family emoji, and a letter with a combining accent each count as one. Code points and UTF-16 length are shown in separate rows, so you can compare them with the length your program reports.
Why is the UTF-16 length different from the character count?
The length of a string in JavaScript and Java is the number of UTF-16 code units. Characters outside the Basic Multilingual Plane, such as most emoji, take two code units, and one visible character can be made of several code points. The thumbs-up emoji with a skin tone is 1 character, 2 code points, and 4 UTF-16 code units.
How do I check text against a length limit?
Enter the limit in the Character limit option. The number of characters left is shown above the table, and when you go over, the excess is highlighted. For a limit in bytes, such as a database column, read the byte row for the encoding that system uses.
What does a dash in a byte row mean?
The text contains characters that the encoding cannot represent. Shift_JIS has no Hangul, for example, and neither EUC-KR nor Shift_JIS has emoji. The number of such characters is shown under the row and in the note below the table.
Related tools
How it works · Standards
- Processed in: this browser (your device). Nothing is sent or stored.
- Standards: WHATWG Encoding · UAX #29
- Engine: ZEKILO Dev character and byte counter (in-house code) · EUC-KR and Shift_JIS encoding tables (built from the WHATWG Encoding indexes) (BSD 3-Clause (WHATWG))
- Characters are segmented with the browser's Intl.Segmenter. In a browser without it, characters are counted as code points and the tool says so.
- The number of characters that cannot be represented is counted in code points.
- Words are pieces separated by whitespace, so a sentence written without spaces counts as one word.
- As of 2026.10.03
- Changelog: 2026.10.03 First release