Skip to main content

HTML Entities

Escape and unescape HTML entities in named, decimal and hex forms

Choose between named, decimal and hex entity formsBy default escapes only the five required characters, keeping output readableLenient decoding: misspelled entities are preserved verbatim and listedIncludes a common-entity reference with characters and code points

Encoding options

Chinese and emoji become entities too, producing pure ASCII output. Only needed for legacy systems

Result9 entity/entities

Common entities

EntityCharacterCode point
&&U+0026
&lt;<U+003C
&gt;>U+003E
&quot;"U+0022
&#39;'U+0027
&nbsp; U+00A0
&copy;©U+00A9
&reg;®U+00AE
&trade;™U+2122
&mdash;—U+2014
&hellip;…U+2026
&times;×U+00D7
&deg;°U+00B0
&euro;€U+20AC
&larr;←U+2190
&rarr;→U+2192

Last updated: 2026-10-10

About this tool

The most common use of HTML entities is to stop user input being executed as markup - in other words, preventing XSS. But escaping every character into a form like &#x4F60; is both unnecessary and unreadable; in practice only five characters must be escaped. This tool lets you choose freely between minimal and complete escaping, and handles misspelled entities leniently when decoding.

Features

Minimal escaping mode

By default only & < > " and ' are escaped - the minimum set required for safety. The result stays readable rather than turning into the unreadable soup of full escaping.

Escape all non-ASCII

When pure ASCII output is required, such as for a legacy system or an email template that only accepts ASCII, enable this to turn Chinese characters and emoji into entities as well.

Three entity forms

Named entities (&amp;), decimal (&#38;) and hex (&#x26;) are all available. They are equivalent; conventions differ by context, and hex is more common in CSS and SVG.

Lenient decoding

Pasted text often contains misspelled entities (&#xZZ;, or named entities that do not exist). These are kept verbatim and listed separately rather than failing the whole decode or silently dropping characters.

Code-point handling

Emoji and rare CJK characters are handled as whole Unicode code points - 😀 correctly becomes &#128512; or &#x1F600; rather than being split into two invalid surrogates.

XSS risk warning

If the input contains unescaped HTML tags the tool warns clearly - inserting that straight into a page is an XSS hole.

How to use

  1. 1

    Choose a direction

    Use “Encode” to write markup safely into a page; use “Decode” to read escaped text from page source back into its original characters.

  2. 2

    Adjust options as needed

    The default minimal escaping suits most cases. Only enable “Escape all non-ASCII” when you specifically need pure ASCII output.

  3. 3

    Paste your text

    Results update live with no button to press. In decode mode, anything listed below the result is a fragment that could not be recognised.

  4. 4

    Copy the result

    Copy it straight into your template or code. If an XSS warning appears, make sure the content is from a trusted source - otherwise escape it first.

Options

Required escapes
Five in total: & (the entity introducer), < and > (tag boundaries), and " and ' (attribute value delimiters). Only these five cause parse errors or security problems when left alone; everything else can stay as-is.
Named entities
Forms like &amp;, &lt; and &copy;, written with a readable name. The advantage is clarity; the drawback is that they are limited - Chinese characters and emoji have no named entities and fall back to numeric form.
Decimal entities
Forms like &#20320;, using a decimal code point to express any character. The most compatible, supported by every browser and parser.
Hex entities
Forms like &#x4F60;, using a hex code point. More common in CSS, SVG and JavaScript, where hexadecimal code points are the convention.
Invalid code points
Code points beyond the Unicode range (above U+10FFFF) or inside the surrogate range (U+D800–U+DFFF) are not legal characters; this tool refuses to decode them and keeps them verbatim.

Common use cases

  • Displaying user input safely on a page (preventing XSS)
  • Turning escaped HTML source back into readable content
  • Producing pure ASCII text output for a legacy system
  • Inserting special symbols such as ©, ™ and → into HTML
  • Finding out which characters actually lie behind a piece of entity text
  • Debugging pages with garbled text or misplaced punctuation
  • Checking whether a particular entity is spelled correctly

FAQ

Questions you may have about this tool

When must HTML entities be used?

Whenever untrusted content is inserted into HTML. The most common cases are & (otherwise AT&T is read as the start of an entity), < and > (otherwise the text is parsed as markup, the main XSS entry point), and quotes (otherwise an attribute value closes early). For every other character, escaping affects readability rather than safety.

Are encoding and escaping the same thing?

In an HTML context they are effectively the same - replacing special characters with entity forms. But note that entities are **not** a defence against SQL injection; that requires parameterised queries. Conflating the two is a dangerous practice.

Named or numeric entities - which should I use?

Functionally equivalent; it depends on context. Named entities (&amp;) are readable and suit hand-written templates; numeric ones (&#38; or &#x26;) express any character and suit programmatic generation. Remember that Chinese characters and emoji have no named entities and are numeric only.

Why does &nbsp; look like a normal space but is not one?

&nbsp; is a non-breaking space (U+00A0). Its very purpose is to **prevent** a line break at that point, and it is rendered slightly wider than a normal space. Because it looks identical, it is a common source of layout bugs - it is often copied in from web pages, and this tool’s reference table lets you confirm the code point is 160 rather than 32.

Will escaped text still render correctly in a browser?

Yes. Entities are part of the HTML specification; the browser decodes them into their characters before handing them to the rendering engine. To the browser, &lt;div&gt; is the text <div> and is never treated as a tag.

What happens to a misspelled entity when decoding?

This tool keeps any fragment it cannot recognise **verbatim** and lists it separately below the result, rather than discarding it or replacing it with a question mark. That way you can see immediately where the problem is instead of hunting character by character for why content went missing.

Why did Chinese characters turn into numbers after enabling full escaping?

Because Chinese has no named entities - the HTML specification defines only about 2,000 named entities, mostly covering Western European characters and mathematical symbols. Chinese, Japanese, Korean and emoji are not among them and fall back to numeric entities. That is a limitation of the specification, not of the tool.

Is my input uploaded?

No. Encoding and decoding are performed entirely locally by your browser with no network requests. Even template content containing sensitive information never leaves your device.