Docento.app
Clean workspace with laptop and notebook
All Posts

Text File Encodings Explained: UTF-8, UTF-16, ANSI and the BOM

By The Docento.app TeamPublished 5 min read
Try Docento's free PDF editor — No sign-up, 100% private — sign, annotate, and stamp PDFs in your browser.Open the editor

A text file is just bytes, and an encoding is the agreement about which characters those bytes stand for. When a program assumes a different encoding from the one used to write the file, accented letters, symbols and emoji turn into strange characters. Understanding the main encodings makes those problems easy to diagnose.

Why encodings exist

Early computers used one byte per character, giving 256 possible characters, which is not enough for every language. Different regions defined different tables, called code pages, and the same byte meant different characters in each. Unicode later assigned a number, called a code point, to every character in every script, and encodings such as UTF-8 and UTF-16 describe how to store those numbers as bytes.

The encodings you will meet

ASCII covers 128 characters: English letters, digits and punctuation, plus control codes. It is a subset of most other encodings, so pure-ASCII files look right nearly everywhere.

UTF-8 encodes Unicode using one to four bytes per character. The first 128 characters are identical to ASCII, so ASCII text is valid UTF-8, which is a big reason it is the dominant encoding on the web, in source code and in most modern tools. It is compact for English text and handles every script.

UTF-16 uses two bytes for most characters and four for others. It is used internally by Windows and by some programming platforms, and is the "Unicode" option in some Windows editors. Because it uses two bytes per character, an ASCII character is followed by a zero byte, and the endianness (byte order) matters.

Windows-1252, often labelled "ANSI" in Windows tools, is a single-byte encoding for Western European languages. ISO-8859-1 (Latin-1) is similar. Other regions have their own code pages, such as Windows-1251 for Cyrillic and Shift JIS for Japanese. "ANSI" is a misleading name, since it just means whatever legacy code page the system uses, so what it covers depends on the machine.

The byte order mark (BOM)

A BOM is the character U+FEFF placed at the start of a file to signal how it is encoded. In UTF-16 it also tells you the byte order. In UTF-8 byte order is not an issue, but some programs write the three bytes EF BB BF anyway as a signature.

  • A UTF-8 BOM helps some Windows programs recognise a file as UTF-8, notably older versions of Excel opening CSV files.
  • It can hurt elsewhere. Parsers for JSON, YAML, .env files and shell scripts may treat it as part of the first token. Unix tools may show a strange character at the start of the first line, and the first key in a config file may not match.

The Unicode standard does not require or recommend a BOM for UTF-8. If your files are consumed by software rather than people with older tools, save as UTF-8 without BOM.

Recognising the symptoms

What you see Likely cause
é instead of é UTF-8 file read as Windows-1252 or Latin-1
� (replacement character) Bytes that are not valid in the encoding being used
Letters separated by spaces or nulls UTF-16 file read as a single-byte encoding
 at the start A UTF-8 BOM shown by a tool that read it as Latin-1
? where letters should be Characters not representable in the target encoding were replaced when saving
Cyrillic or Greek letters shown as accented Latin ones Wrong code page selected

How to fix it

  1. Do not save over the file while it looks wrong. Saving a misread file can make the damage permanent.
  2. Open with the correct encoding. Many editors let you reopen a file choosing the encoding. Try UTF-8 first.
  3. Convert once you know. Re-save as UTF-8 for new work.
  4. For data files, use the import dialog's encoding option instead of double-clicking. See fixing garbled characters in CSV files.

How Docento handles encodings

The browser version of Docento's Text & Markdown Editor reads files as UTF-8 and saves what you download as UTF-8, without a byte order mark. If your file shows mojibake there, it is probably in another encoding, so convert it first or open it in an editor that lets you choose.

Prevention

  • Use UTF-8 everywhere by default.
  • Declare the encoding where the format allows it, such as in an XML declaration or an HTTP header.
  • Be cautious with files that pass through spreadsheets and word processors.
  • Test with accented letters and an emoji before relying on a workflow.

Takeaway

Use UTF-8 without a BOM as the default, recognise the telltale symptoms of a mismatch and always reopen with the right encoding before you save. Saving a misread file is what turns a display problem into data loss.

Try Docento's free PDF editor

No sign-up, 100% private — sign, annotate, and stamp PDFs in your browser.

Open the editor

Related Posts