HTML Entities
Escape and unescape HTML entities in named, decimal and hex forms
Encoding options
Chinese and emoji become entities too, producing pure ASCII output. Only needed for legacy systems
Unescaped HTML tags detected. If this content comes from user input, inserting it into a page opens an XSS hole - escape it first.
Result9 entity/entities
Common entities
| Entity | Character | Code point |
|---|---|---|
| & | & | U+0026 |
| < | < | U+003C |
| > | > | U+003E |
| " | " | U+0022 |
| ' | ' | U+0027 |
| | U+00A0 | |
| © | © | U+00A9 |
| ® | ® | U+00AE |
| ™ | ™ | U+2122 |
| — | — | U+2014 |
| … | … | U+2026 |
| × | × | U+00D7 |
| ° | ° | U+00B0 |
| € | € | U+20AC |
| ← | ← | U+2190 |
| → | → | U+2192 |
Last updated: 2026-10-10
About this tool
The most common use of HTML entities is to stop user input being executed as markup - in other words, preventing XSS. But escaping every character into a form like 你 is both unnecessary and unreadable; in practice only five characters must be escaped. This tool lets you choose freely between minimal and complete escaping, and handles misspelled entities leniently when decoding.
Features
Minimal escaping mode
By default only & < > " and ' are escaped - the minimum set required for safety. The result stays readable rather than turning into the unreadable soup of full escaping.
Escape all non-ASCII
When pure ASCII output is required, such as for a legacy system or an email template that only accepts ASCII, enable this to turn Chinese characters and emoji into entities as well.
Three entity forms
Named entities (&), decimal (&) and hex (&) are all available. They are equivalent; conventions differ by context, and hex is more common in CSS and SVG.
Lenient decoding
Pasted text often contains misspelled entities (&#xZZ;, or named entities that do not exist). These are kept verbatim and listed separately rather than failing the whole decode or silently dropping characters.
Code-point handling
Emoji and rare CJK characters are handled as whole Unicode code points - 😀 correctly becomes 😀 or 😀 rather than being split into two invalid surrogates.
XSS risk warning
If the input contains unescaped HTML tags the tool warns clearly - inserting that straight into a page is an XSS hole.
How to use
- 1
Choose a direction
Use “Encode” to write markup safely into a page; use “Decode” to read escaped text from page source back into its original characters.
- 2
Adjust options as needed
The default minimal escaping suits most cases. Only enable “Escape all non-ASCII” when you specifically need pure ASCII output.
- 3
Paste your text
Results update live with no button to press. In decode mode, anything listed below the result is a fragment that could not be recognised.
- 4
Copy the result
Copy it straight into your template or code. If an XSS warning appears, make sure the content is from a trusted source - otherwise escape it first.
Options
- Required escapes
- Five in total: & (the entity introducer), < and > (tag boundaries), and " and ' (attribute value delimiters). Only these five cause parse errors or security problems when left alone; everything else can stay as-is.
- Named entities
- Forms like &, < and ©, written with a readable name. The advantage is clarity; the drawback is that they are limited - Chinese characters and emoji have no named entities and fall back to numeric form.
- Decimal entities
- Forms like 你, using a decimal code point to express any character. The most compatible, supported by every browser and parser.
- Hex entities
- Forms like 你, using a hex code point. More common in CSS, SVG and JavaScript, where hexadecimal code points are the convention.
- Invalid code points
- Code points beyond the Unicode range (above U+10FFFF) or inside the surrogate range (U+D800–U+DFFF) are not legal characters; this tool refuses to decode them and keeps them verbatim.
Common use cases
- Displaying user input safely on a page (preventing XSS)
- Turning escaped HTML source back into readable content
- Producing pure ASCII text output for a legacy system
- Inserting special symbols such as ©, ™ and → into HTML
- Finding out which characters actually lie behind a piece of entity text
- Debugging pages with garbled text or misplaced punctuation
- Checking whether a particular entity is spelled correctly
FAQ
Questions you may have about this tool
When must HTML entities be used?
Whenever untrusted content is inserted into HTML. The most common cases are & (otherwise AT&T is read as the start of an entity), < and > (otherwise the text is parsed as markup, the main XSS entry point), and quotes (otherwise an attribute value closes early). For every other character, escaping affects readability rather than safety.
Are encoding and escaping the same thing?
In an HTML context they are effectively the same - replacing special characters with entity forms. But note that entities are **not** a defence against SQL injection; that requires parameterised queries. Conflating the two is a dangerous practice.
Named or numeric entities - which should I use?
Functionally equivalent; it depends on context. Named entities (&) are readable and suit hand-written templates; numeric ones (& or &) express any character and suit programmatic generation. Remember that Chinese characters and emoji have no named entities and are numeric only.
Why does look like a normal space but is not one?
is a non-breaking space (U+00A0). Its very purpose is to **prevent** a line break at that point, and it is rendered slightly wider than a normal space. Because it looks identical, it is a common source of layout bugs - it is often copied in from web pages, and this tool’s reference table lets you confirm the code point is 160 rather than 32.
Will escaped text still render correctly in a browser?
Yes. Entities are part of the HTML specification; the browser decodes them into their characters before handing them to the rendering engine. To the browser, <div> is the text <div> and is never treated as a tag.
What happens to a misspelled entity when decoding?
This tool keeps any fragment it cannot recognise **verbatim** and lists it separately below the result, rather than discarding it or replacing it with a question mark. That way you can see immediately where the problem is instead of hunting character by character for why content went missing.
Why did Chinese characters turn into numbers after enabling full escaping?
Because Chinese has no named entities - the HTML specification defines only about 2,000 named entities, mostly covering Western European characters and mathematical symbols. Chinese, Japanese, Korean and emoji are not among them and fall back to numeric entities. That is a limitation of the specification, not of the tool.
Is my input uploaded?
No. Encoding and decoding are performed entirely locally by your browser with no network requests. Even template content containing sensitive information never leaves your device.