Skip to main content
ZEKILO Dev
Encode

Why Non-ASCII URLs Turn Into %EC%A0%9C: Percent-Encoding

Updated 2026.10.03 · 3 min read

URLs allow only part of ASCII, so other characters are written as %XX bytes. See why UTF-8 and EUC-KR results differ and when to use encodeURIComponent.

Open URL Encode & Decode

When a link with Korean, Japanese or other non-ASCII text turns into something like %EC%A0%9C…, it is not corrupted. It is percent-encoded. A URL may contain only part of ASCII, so any other character is first converted to bytes, normally in UTF-8, and each byte is written as % followed by two hexadecimal digits. The Korean syllable 제 is three bytes in UTF-8, EC A0 9C, so it becomes %EC%A0%9C. Switch URL Encode & Decode to Decode and paste the link to get the original characters back; set the character encoding to EUC-KR to read links that were encoded that way.

The rules of percent-encoding

Following the rules of RFC 3986, the characters of a URL fall into three groups.

Group Characters Handling
Unreserved characters Letters, digits, - . _ ~ Written as they are
Reserved characters :/?#[]@ and !$&'()*+,;= Act as delimiters; encode them when used as data
Everything else Non-ASCII characters, spaces and some ASCII characters such as "<> Converted to bytes and written as %XX

Uppercase hexadecimal digits are recommended, and %ec and %EC mean the same byte.

Input Result (UTF-8, encodeURIComponent)
제킬로 %EC%A0%9C%ED%82%AC%EB%A1%9C
제킬로 dev&tools=1 %EC%A0%9C%ED%82%AC%EB%A1%9C%20dev%26tools%3D1

One Hangul syllable is three bytes in UTF-8, so it grows to nine characters once encoded. That is why links with non-ASCII text get so long.

UTF-8 and legacy encodings give different results

Percent-encoding is only a way to write bytes. Which bytes a character becomes is decided by the character encoding, so the same word 제킬로 has two different results.

Character encoding Bytes Percent-encoded
UTF-8 EC A0 9C ED 82 AC EB A1 9C (9) %EC%A0%9C%ED%82%AC%EB%A1%9C
EUC-KR (CP949) C1 A6 C5 B3 B7 CE (6) %C1%A6%C5%B3%B7%CE

RFC 3986 recommends that new URI schemes convert character data to UTF-8 before percent-encoding it. The WHATWG URL Standard, which browsers follow, always encodes the path in UTF-8. The query string, however, follows the character encoding of the document, so links and forms on a page written in EUC-KR produce values like %C1%A6…. The same thing happens with Shift_JIS on older Japanese pages. This is why an integration with an older system may ask you to encode parameters in a legacy encoding.

Some ways to tell the two apart:

  • If one Hangul syllable is three %XX groups, the text is probably UTF-8; two groups suggest EUC-KR.
  • Decoding an EUC-KR result as UTF-8 fails. In JavaScript, decodeURIComponent('%C1%A6%C5%B3%B7%CE') throws a URIError.
  • Decoding a UTF-8 result as EUC-KR either fails or gives unreadable characters.

Do not guess which encoding the other system uses. Check its integration documentation.

encodeURIComponent, encodeURI and form encoding

Method Characters left as they are Space Use it for
encodeURIComponent Letters, digits, -_.!~*'() %20 One query value or path segment
encodeURI The above plus ;/?:@&=+$,# %20 A whole URL
Form encoding Letters, digits, *-._ + HTML forms, URLSearchParams
  • Whole URL: https://dev.zekilo.com/ko/검색?q=a b becomes https://dev.zekilo.com/ko/%EA%B2%80%EC%83%89?q=a%20b. The delimiters :, /, ? and = stay.
  • Form encoding: a b&c=제 becomes a+b%26c%3D%EC%A0%9C. The space turns into +.
  • !'()*~ is not changed by encodeURIComponent.

If a value contains & or = and you encode it with encodeURI, those characters are read as delimiters and the parameter is split. For a single query value, use encodeURIComponent or URLSearchParams.

Problems you are likely to meet

  • Double encoding: encoding a string that is already encoded turns each % into %25, so %EC%A0%9C becomes %25EC%25A0%259C. RFC 3986 says the same string must not be percent-encoded or decoded more than once.
  • Truncated sequences: if a link is cut off and ends like %E0%A4%A, decoding fails. Check whether a long link lost its end in a chat message or an email.
  • + and spaces: only form decoding reads + as a space. decodeURIComponent('a+b') returns a+b unchanged.
  • Non-ASCII domain names: the host name is not percent-encoded; it is converted to Punycode. The host of https://한글.kr/ becomes xn--bj0bj06e.kr.

Summary

  • %EC%A0%9C is simply the UTF-8 bytes of 제 written out. Decoding gives the original character back.
  • Use UTF-8 in anything new. Use EUC-KR or Shift_JIS only when the other system requires it.
  • Use encodeURIComponent for values and encodeURI for whole URLs, and encode only once.
  • To inspect a link, use URL Encode & Decode. Your input is processed only in your browser.

Tools for this guide

Sources

More guides