Characters vs Bytes: UTF-8, EUC-KR and Shift_JIS
Updated 2026.10.03 · 4 min read
Text has different byte counts in UTF-8, UTF-16, EUC-KR (CP949) and Shift_JIS (CP932). See the numbers and how to fit VARCHAR and fixed-length records.
Open Character & Byte CounterA character count and a byte count are different units. 안녕하세요 ZEKILO is 12 characters, but 22 bytes in UTF-8 and 17 bytes in EUC-KR. Wherever a length is measured in bytes, such as a database column or a fixed-length record, you have to count in the character encoding the other system uses. Put your text into the Character & Byte Counter to see the character count, the UTF-16 length and the UTF-8, EUC-KR and Shift_JIS byte counts side by side.
How many bytes is one character?
All numbers in this table are bytes.
| Character | UTF-8 | UTF-16 | EUC-KR (CP949) | Shift_JIS (CP932) |
|---|---|---|---|---|
A letter or digit |
1 | 2 | 1 | 1 |
가 Hangul |
3 | 2 | 2 | Not representable |
あ hiragana |
3 | 2 | 2 | 2 |
漢 kanji |
3 | 2 | 2 | 2 |
ア half-width katakana |
3 | 2 | Not representable | 1 |
😀 emoji |
4 | 4 | Not representable | Not representable |
- UTF-8 uses one to four bytes depending on the code point (RFC 3629): one byte up to U+007F, two up to U+07FF, three up to U+FFFF and four above that. Hangul, kana and common kanji take three bytes.
- UTF-16 works in 2-byte units. A character above U+FFFF takes two units (four bytes). JavaScript’s
lengthand Java’sString.length()return the number of these units, not bytes. - EUC-KR writes ASCII in one byte and Hangul, kana and kanji in two. Shift_JIS writes ASCII and half-width katakana in one byte and kana and kanji in two; it cannot write Hangul.
EUC-KR follows CP949 (MS949) and Shift_JIS follows CP932 (Windows-31J). Other systems may differ. In the WHATWG Encoding Standard, euc-kr means KS X 1001 together with the Unified Hangul Code, in other words Windows code page 949.
The difference in examples
| Input | Characters | UTF-16 length | UTF-8 | EUC-KR (CP949) | Shift_JIS (CP932) |
|---|---|---|---|---|---|
안녕하세요 ZEKILO |
12 | 12 | 22 | 17 | 5 Hangul are not representable |
こんにちは ZEKILO |
12 | 12 | 22 | 17 | 17 |
👍🏻가 |
2 | 5 | 11 | — | — |
똠방각하 |
4 | 4 | 12 | 8 | — |
①髙 |
2 | 2 | 6 | — | 4 |
A dash (—) means the input contains characters that the encoding cannot represent.
- The thumbs-up in
👍🏻가is made of two code points, the emoji and a skin tone, but it is displayed as one character. Counted the way a person sees it (grapheme clusters, UAX #29) the input is 2 characters; it is 3 code points and its UTF-16 length is 5. The answer depends on what you count. 똠is not among the 2,350 Hangul syllables of KS X 1001. CP949 has it in its extension area and writes it in two bytes, but a system that supports only the KS X 1001 range may not be able to store it.①and髙are not in JIS X 0208; they are CP932 extensions.
Database VARCHAR: is n characters or bytes?
What the 10 in VARCHAR(10) counts depends on the product (according to each product’s documentation as of October 2026).
| Product | Unit of n |
|---|---|
MySQL VARCHAR(n) |
Characters |
PostgreSQL varchar(n) |
Characters |
Oracle VARCHAR2(n) |
BYTE or CHAR as declared; if omitted, NLS_LENGTH_SEMANTICS (default BYTE) applies |
SQL Server varchar(n) |
Bytes |
In a byte-based column the character encoding decides the limit. The Korean name 홍길동 is 6 bytes in EUC-KR and 9 bytes in UTF-8. A column of 10 bytes holds five Hangul syllables in EUC-KR but only three in UTF-8. This is why moving a database from EUC-KR to UTF-8 causes length errors: every Hangul syllable grows from two bytes to three. Before migrating, count the bytes of the longest values in the new encoding.
Even MySQL, which counts characters, has a limit of 65,535 bytes for the whole row, so the maximum length you can actually declare depends on the character set.
Fixed-length records
In a fixed-length record every field has a set number of bytes, and unused positions are filled with spaces or zeros. If one field is off by a single byte, every field after it shifts.
- Calculate padding in bytes. Putting
홍길동into a 10-byte name field in EUC-KR takes 6 bytes, so you add 4 spaces. If you calculate from the character count (3) and add 7 spaces, the field becomes 13 bytes. - Cut at a character boundary. If only the first byte of a two-byte character is left, that character and what follows it are damaged. Add characters one at a time and stop before the limit would be exceeded.
- A line break is 2 bytes as CRLF and 1 byte as LF.
- Filter out characters the other system cannot represent, such as emoji or kanji that are not in CP949. The Character & Byte Counter warns you how many characters cannot be represented in EUC-KR or Shift_JIS.
Summary
- When you see a length limit, first find out whether the unit is characters or bytes, and if it is bytes, in which character encoding.
- Hangul is 3 bytes in UTF-8 and 2 bytes in EUC-KR. An emoji such as
😀is 4 bytes in UTF-8 and cannot be written in EUC-KR or Shift_JIS. - The value returned by
lengthis the UTF-16 length, which can differ from both the character count and the byte count. - Before you send data, check the real byte count with the Character & Byte Counter. Your input is processed only in your browser.
Tools for this guide
Sources
- WHATWG Encoding Standard
- RFC 3629 — UTF-8, a transformation format of ISO 10646
- Unicode UAX #29 — Unicode Text Segmentation
- MySQL 8.4 Reference Manual — The CHAR and VARCHAR Types
- PostgreSQL Documentation — Character Types
- Oracle Database 19c SQL Language Reference — Data Types
- Oracle Database 19c Reference — NLS_LENGTH_SEMANTICS
- Microsoft Learn — char and varchar (Transact-SQL)