Many XML errors come down to five characters. Because XML uses < and & as syntax, those characters cannot simply appear in text. Understanding how to escape them solves the most common "not well-formed" failures.
The five predefined entities
XML defines five built-in entity references:
| Character | Entity |
|---|---|
< |
< |
> |
> |
& |
& |
" |
" |
' |
' |
Which ones you must use depends on where the character appears:
- In text content,
<and&must always be escaped.>should be escaped when it appears in the sequence]]>, and escaping it is conventional. - In attribute values, escape
<and&, plus whichever quote character delimits the value.
<company name="Smith & Sons">AT&T <tag></company>
The classic error
<url>https://example.com/?a=1&b=2</url>
The &b=2 looks like the start of an entity reference, and the parser finds no semicolon. The result is an error along the lines of "EntityRef: expecting ';'" or "unescaped ampersand". The correct text is a=1&b=2. This happens constantly with URLs and company names.
Character references
You can write any Unicode character by its code point:
- Decimal:
©is © - Hexadecimal:
©is ©
This is useful for characters your file's encoding cannot represent. If the file is UTF-8, you can usually type the character directly.
CDATA sections
When a block contains lots of special characters, such as code or markup that should be treated as plain text, a CDATA section avoids escaping:
<script><![CDATA[
if (a < b && b > c) { run(); }
]]></script>
Inside CDATA, < and & are literal text. The only thing it cannot contain is the sequence ]]>, which ends the section. If you need that exact sequence, split it across two CDATA sections. CDATA works in element content only, not in attribute values, and it does not exist in the processed data: parsers hand you the same text either way.
Other character problems
- Illegal control characters. XML 1.0 forbids most control characters (below U+0020, other than tab, line feed and carriage return), even when escaped as character references. Text pasted from terminals or databases sometimes contains them. Strip them before writing XML.
- Encoding mismatches. The declaration
encoding="UTF-8"must match the real encoding. See text file encodings. - Non-breaking spaces and smart quotes are fine as content, but smart quotes cannot delimit attribute values.
- Custom entities.
is an HTML entity, not an XML one, and causes an "undefined entity" error unless a DTD defines it. Use instead.
Comments
Comments cannot contain the double hyphen --, and cannot end in -, and a comment inside an element must not be placed inside a tag. These restrictions surprise people who comment out large blocks that already contain comments.
How to escape in practice
- Generate XML with a library, not by joining strings. Libraries escape correctly.
- When editing by hand, replace
&first, then<and>, to avoid double-escaping. - Use CDATA for code blocks.
- Check the result. Docento's Text & Markdown Editor parses
.xmlfiles in the browser, and a stray&or<will show up immediately with its line number. See how to open and read XML files.
A quick checklist
- Does text contain
&or<? Escape them. - Is the attribute value quoted, and are the matching quotes escaped inside it?
- Are there HTML-only entities like
? Replace them. - Does the file contain control characters? Remove them.
- Does the declared encoding match the file?
Takeaway
Escape & and < in XML text, escape the delimiting quote in attributes, use numeric references for HTML-only characters and CDATA for blocks of code. Better still, let an XML library do the escaping for you.