HTML Entity Encoder and Decoder

Write ampersands and angle brackets as character references, or turn known references back into characters.

Inputs stay on your device No sign-up Free to use
How this works

The tool runs in this browser. Your file or text is not uploaded to UseFreeTools. Check this tool's limits for anything it may save on your device.

Privacy details

Convert the entities controls

Showing an example. Edit to see your own.

Paste the text to encode, or the text containing references to decode. The work happens in this browser.

Used when encoding. Decoding reads named, decimal and hexadecimal references whatever style is chosen here.

Used by the two numeric styles. The named style always encodes them, because it has no letter to leave in place.

Used when decoding. A missing semicolon, an unlisted name and an impossible code point are reported with their position.

Processed in your browser. Your inputs stay on this device.

How to use HTML Entity Encoder and Decoder

  1. Choose the Direction: text to references, or references to text.
  2. Paste the text in Text.
  3. For encoding, pick a Reference style: named, decimal or hexadecimal.
  4. The converted text updates as you change the fields. Copy the result or download the .txt file.

Example: HTML Entity Encoder and Decoder

Decode a single escaped reference.

You add
Direction: References to text. Text: <
You get
The summary reports Decoded 1 references once, so the output is &lt; rather than <. The downloaded decoded-text.txt holds the same one-line result.

Options

Reference style
Named writes a name where one exists, such as &amp; and &copy;, and a decimal number for the rest. The numeric styles write numbers, which are easier to scan in a log.
Also encode accented letters and other non-ASCII characters
Used by the two numeric styles. Turning it on writes references for accented letters too. The named style already numbers any character it has no name for.
When a reference is unknown or malformed
Used when decoding. Leave it keeps the source as typed and lists the problems. Stop and report the first one suits a typo hunt.

Supported inputs and limits

Up to 200,000 characters in one run. Names cover 34 common references; every other character the named style writes becomes a decimal number. Decoding reads named, decimal and hexadecimal references whatever style you chose. Numeric references are checked first, so a value outside the Unicode range, a surrogate, a null or a control character is reported instead of written. The C1 range follows the HTML standard, which maps it through Windows-1252, so &#151; decodes to an em dash. This is not a sanitizer: it gives you text, not a decision about where that text may go.

Where your input is processed

This tool processes your input in this browser. Your text and files are not uploaded to UseFreeTools. Check this tool's limits for anything it may save on your device.

Where the reference names come from

Named references are fixed by the HTML standard, which holds far more names than the 34 this tool writes. The encoder keeps the ones people meet every day: the five characters that change meaning in markup, plus common punctuation, currency, fractions and arrows. Anything outside that list becomes a number instead, which is longer but always resolves.

WHATWG, Named character references

Questions about HTML Entity Encoder and Decoder

Is the decoded text safe to put straight into my page?

Decoding gives you characters. Whether they are safe depends on where they land, since element content, an attribute and a script block each have their own rules. Escape for that spot.

Why does decoding &amp;lt; give me &lt; instead of a less-than sign?

One pass removes one layer, so the outer reference became an ampersand and the letters lt. Run the decoder again if you want that layer, but check the value first.

Why is &#151; shown as an em dash?

The HTML standard does not read that range as control codes. It maps those numbers through Windows-1252 first, so 151 becomes an em dash.

What happens to control characters?

Named style leaves them in the text as they are rather than writing a reference. On decode, a numeric reference that points at a control character, a null or a surrogate is reported with its position.

Project manager: Tony Hines · Content updated 29 September 2026 · Report a problem