Skip to main content
ZEKILO Dev
Text

Characters vs Bytes: UTF-8, EUC-KR and Shift_JIS

Updated 2026.10.03 · 4 min read

Text has different byte counts in UTF-8, UTF-16, EUC-KR (CP949) and Shift_JIS (CP932). See the numbers and how to fit VARCHAR and fixed-length records.

Open Character & Byte Counter

A character count and a byte count are different units. 안녕하세요 ZEKILO is 12 characters, but 22 bytes in UTF-8 and 17 bytes in EUC-KR. Wherever a length is measured in bytes, such as a database column or a fixed-length record, you have to count in the character encoding the other system uses. Put your text into the Character & Byte Counter to see the character count, the UTF-16 length and the UTF-8, EUC-KR and Shift_JIS byte counts side by side.

How many bytes is one character?

All numbers in this table are bytes.

Character UTF-8 UTF-16 EUC-KR (CP949) Shift_JIS (CP932)
A letter or digit 1 2 1 1
가 Hangul 3 2 2 Not representable
あ hiragana 3 2 2 2
漢 kanji 3 2 2 2
ア half-width katakana 3 2 Not representable 1
😀 emoji 4 4 Not representable Not representable
  • UTF-8 uses one to four bytes depending on the code point (RFC 3629): one byte up to U+007F, two up to U+07FF, three up to U+FFFF and four above that. Hangul, kana and common kanji take three bytes.
  • UTF-16 works in 2-byte units. A character above U+FFFF takes two units (four bytes). JavaScript’s length and Java’s String.length() return the number of these units, not bytes.
  • EUC-KR writes ASCII in one byte and Hangul, kana and kanji in two. Shift_JIS writes ASCII and half-width katakana in one byte and kana and kanji in two; it cannot write Hangul.

EUC-KR follows CP949 (MS949) and Shift_JIS follows CP932 (Windows-31J). Other systems may differ. In the WHATWG Encoding Standard, euc-kr means KS X 1001 together with the Unified Hangul Code, in other words Windows code page 949.

The difference in examples

Input Characters UTF-16 length UTF-8 EUC-KR (CP949) Shift_JIS (CP932)
안녕하세요 ZEKILO 12 12 22 17 5 Hangul are not representable
こんにちは ZEKILO 12 12 22 17 17
👍🏻가 2 5 11 — —
똠방각하 4 4 12 8 —
①髙 2 2 6 — 4

A dash (—) means the input contains characters that the encoding cannot represent.

  • The thumbs-up in 👍🏻가 is made of two code points, the emoji and a skin tone, but it is displayed as one character. Counted the way a person sees it (grapheme clusters, UAX #29) the input is 2 characters; it is 3 code points and its UTF-16 length is 5. The answer depends on what you count.
  • 똠 is not among the 2,350 Hangul syllables of KS X 1001. CP949 has it in its extension area and writes it in two bytes, but a system that supports only the KS X 1001 range may not be able to store it.
  • ① and 髙 are not in JIS X 0208; they are CP932 extensions.

Database VARCHAR: is n characters or bytes?

What the 10 in VARCHAR(10) counts depends on the product (according to each product’s documentation as of October 2026).

Product Unit of n
MySQL VARCHAR(n) Characters
PostgreSQL varchar(n) Characters
Oracle VARCHAR2(n) BYTE or CHAR as declared; if omitted, NLS_LENGTH_SEMANTICS (default BYTE) applies
SQL Server varchar(n) Bytes

In a byte-based column the character encoding decides the limit. The Korean name 홍길동 is 6 bytes in EUC-KR and 9 bytes in UTF-8. A column of 10 bytes holds five Hangul syllables in EUC-KR but only three in UTF-8. This is why moving a database from EUC-KR to UTF-8 causes length errors: every Hangul syllable grows from two bytes to three. Before migrating, count the bytes of the longest values in the new encoding.

Even MySQL, which counts characters, has a limit of 65,535 bytes for the whole row, so the maximum length you can actually declare depends on the character set.

Fixed-length records

In a fixed-length record every field has a set number of bytes, and unused positions are filled with spaces or zeros. If one field is off by a single byte, every field after it shifts.

  • Calculate padding in bytes. Putting 홍길동 into a 10-byte name field in EUC-KR takes 6 bytes, so you add 4 spaces. If you calculate from the character count (3) and add 7 spaces, the field becomes 13 bytes.
  • Cut at a character boundary. If only the first byte of a two-byte character is left, that character and what follows it are damaged. Add characters one at a time and stop before the limit would be exceeded.
  • A line break is 2 bytes as CRLF and 1 byte as LF.
  • Filter out characters the other system cannot represent, such as emoji or kanji that are not in CP949. The Character & Byte Counter warns you how many characters cannot be represented in EUC-KR or Shift_JIS.

Summary

  • When you see a length limit, first find out whether the unit is characters or bytes, and if it is bytes, in which character encoding.
  • Hangul is 3 bytes in UTF-8 and 2 bytes in EUC-KR. An emoji such as 😀 is 4 bytes in UTF-8 and cannot be written in EUC-KR or Shift_JIS.
  • The value returned by length is the UTF-16 length, which can differ from both the character count and the byte count.
  • Before you send data, check the real byte count with the Character & Byte Counter. Your input is processed only in your browser.

Tools for this guide

Sources

More guides