Why Non-ASCII URLs Turn Into %EC%A0%9C: Percent-Encoding
Updated 2026.10.03 · 3 min read
URLs allow only part of ASCII, so other characters are written as %XX bytes. See why UTF-8 and EUC-KR results differ and when to use encodeURIComponent.
Open URL Encode & DecodeWhen a link with Korean, Japanese or other non-ASCII text turns into something like %EC%A0%9C…, it is not corrupted. It is percent-encoded. A URL may contain only part of ASCII, so any other character is first converted to bytes, normally in UTF-8, and each byte is written as % followed by two hexadecimal digits. The Korean syllable 제 is three bytes in UTF-8, EC A0 9C, so it becomes %EC%A0%9C. Switch URL Encode & Decode to Decode and paste the link to get the original characters back; set the character encoding to EUC-KR to read links that were encoded that way.
The rules of percent-encoding
Following the rules of RFC 3986, the characters of a URL fall into three groups.
| Group | Characters | Handling |
|---|---|---|
| Unreserved characters | Letters, digits, - . _ ~ |
Written as they are |
| Reserved characters | :/?#[]@ and !$&'()*+,;= |
Act as delimiters; encode them when used as data |
| Everything else | Non-ASCII characters, spaces and some ASCII characters such as "<> |
Converted to bytes and written as %XX |
Uppercase hexadecimal digits are recommended, and %ec and %EC mean the same byte.
| Input | Result (UTF-8, encodeURIComponent) |
|---|---|
제킬로 |
%EC%A0%9C%ED%82%AC%EB%A1%9C |
제킬로 dev&tools=1 |
%EC%A0%9C%ED%82%AC%EB%A1%9C%20dev%26tools%3D1 |
One Hangul syllable is three bytes in UTF-8, so it grows to nine characters once encoded. That is why links with non-ASCII text get so long.
UTF-8 and legacy encodings give different results
Percent-encoding is only a way to write bytes. Which bytes a character becomes is decided by the character encoding, so the same word 제킬로 has two different results.
| Character encoding | Bytes | Percent-encoded |
|---|---|---|
| UTF-8 | EC A0 9C ED 82 AC EB A1 9C (9) |
%EC%A0%9C%ED%82%AC%EB%A1%9C |
| EUC-KR (CP949) | C1 A6 C5 B3 B7 CE (6) |
%C1%A6%C5%B3%B7%CE |
RFC 3986 recommends that new URI schemes convert character data to UTF-8 before percent-encoding it. The WHATWG URL Standard, which browsers follow, always encodes the path in UTF-8. The query string, however, follows the character encoding of the document, so links and forms on a page written in EUC-KR produce values like %C1%A6…. The same thing happens with Shift_JIS on older Japanese pages. This is why an integration with an older system may ask you to encode parameters in a legacy encoding.
Some ways to tell the two apart:
- If one Hangul syllable is three
%XXgroups, the text is probably UTF-8; two groups suggest EUC-KR. - Decoding an EUC-KR result as UTF-8 fails. In JavaScript,
decodeURIComponent('%C1%A6%C5%B3%B7%CE')throws aURIError. - Decoding a UTF-8 result as EUC-KR either fails or gives unreadable characters.
Do not guess which encoding the other system uses. Check its integration documentation.
encodeURIComponent, encodeURI and form encoding
| Method | Characters left as they are | Space | Use it for |
|---|---|---|---|
encodeURIComponent |
Letters, digits, -_.!~*'() |
%20 |
One query value or path segment |
encodeURI |
The above plus ;/?:@&=+$,# |
%20 |
A whole URL |
| Form encoding | Letters, digits, *-._ |
+ |
HTML forms, URLSearchParams |
- Whole URL:
https://dev.zekilo.com/ko/검색?q=a bbecomeshttps://dev.zekilo.com/ko/%EA%B2%80%EC%83%89?q=a%20b. The delimiters:,/,?and=stay. - Form encoding:
a b&c=제becomesa+b%26c%3D%EC%A0%9C. The space turns into+. !'()*~is not changed byencodeURIComponent.
If a value contains & or = and you encode it with encodeURI, those characters are read as delimiters and the parameter is split. For a single query value, use encodeURIComponent or URLSearchParams.
Problems you are likely to meet
- Double encoding: encoding a string that is already encoded turns each
%into%25, so%EC%A0%9Cbecomes%25EC%25A0%259C. RFC 3986 says the same string must not be percent-encoded or decoded more than once. - Truncated sequences: if a link is cut off and ends like
%E0%A4%A, decoding fails. Check whether a long link lost its end in a chat message or an email. +and spaces: only form decoding reads+as a space.decodeURIComponent('a+b')returnsa+bunchanged.- Non-ASCII domain names: the host name is not percent-encoded; it is converted to Punycode. The host of
https://한글.kr/becomesxn--bj0bj06e.kr.
Summary
%EC%A0%9Cis simply the UTF-8 bytes of 제 written out. Decoding gives the original character back.- Use UTF-8 in anything new. Use EUC-KR or Shift_JIS only when the other system requires it.
- Use
encodeURIComponentfor values andencodeURIfor whole URLs, and encode only once. - To inspect a link, use URL Encode & Decode. Your input is processed only in your browser.