We’re moving to a new home! Our website is currently in test mode while we update and transfer our content. Some pages and resources may be temporarily unavailable. We’ll be back with all resources shortly. Thank you for your patience.

CIE IGCSE | 1.2.1 Representing Text

Lesson objective

Explain how character sets map text to numeric codes represented in binary. Compare ASCII and Unicode and understand why shared encoding rules matter.

Learn

1.2.1 | REPRESENTING TEXT

01 | FROM A MESSAGE TO BITS

When you type “Hi”, the computer does not store two tiny pictures of letters as the text itself. It represents the characters using agreed numeric codes, which are stored and processed in binary.

A character set defines a collection of characters and their codes. A character can be a letter, digit, punctuation mark, space or another symbol. Some codes describe control functions rather than printable symbols.

When text is displayed, software interprets the codes and uses a font to draw the characters. The stored character and its visual appearance are related but different things.

02 | A SHARED CODEBOOK

Imagine a class inventing a code where A means 1 and B means 2. A message is useful only if the receiver uses the same rules. Standard character systems let different computers interpret text consistently.

A small extract from ASCII
CharacterASCII denary codeBinary, displayed in a byte
A6501000001
B6601000010
a9701100001
04800110000
Space3200100000

Uppercase A and lowercase a have different codes because they are different characters. A space also has a code even though it leaves no visible ink on the screen.

You do not need to memorise the entire table. Learn how a mapping works and use a supplied table when answering a question.

03 | ASCII: A SMALL STANDARD CHARACTER SET

ASCII stands for American Standard Code for Information Interchange. Standard ASCII uses seven bits, giving 2⁷ = 128 possible codes, numbered 0 to 127.

It includes English letters, digits, punctuation and control codes. It does not cover the world’s writing systems or emoji.

ASCII values are often displayed or stored in an eight-bit byte with a leading 0. That does not change standard ASCII’s seven-bit code range.

Some older eight-bit systems are called extended ASCII. There is not one universal extended-ASCII table; the interpretation of codes above 127 depends on the particular encoding.

04 | ENCODE AND DECODE A SHORT WORD

Encoding CAT using ASCII
CharacterDenary codeBinary byte
C6701000011
A6501000001
T8401010100
Text:    CAT
Codes:   67       65       84
Binary:  01000011 01000001 01010100

Encoding converts characters into a representation for storage or transmission. Decoding interprets that representation back into characters.

To decode, split this example into bytes, convert each byte to a number and find the character associated with that code. The spaces between groups are just for readability here; they are not extra space characters in CAT.

05 | A TEXT DIGIT IS NOT THE NUMERIC VALUE

The character “5” has ASCII code 53. Its displayed binary byte is 00110101. The unsigned integer value 5 can instead be written as 00000101.

Text character "5":  00110101  (ASCII code 53)
Unsigned integer 5:  00000101  (numeric value 5)

The application needs to know what kind of data the bits represent. A string such as “123” contains three characters, not automatically a single binary integer.

Likewise, changing a font does not necessarily change the character codes. A and A in two different fonts can still represent the same character.

06 | UNICODE: MANY LANGUAGES AND SYMBOLS

A global messaging system needs more than English letters. Unicode supports a much greater range of characters and symbols than ASCII, including many writing systems and emoji.

Examples from Unicode
CharacterUnicode code pointWhat it represents
AU+0041Latin capital letter A
éU+00E9Latin small letter e with acute
ΩU+03A9Greek capital letter omega
中U+4E2DA CJK ideograph
😀U+1F600Grinning face emoji

A code point is a number assigned in Unicode. U+ is a notation prefix and the following digits are hexadecimal. A code-point label is not itself the sequence of bytes stored in a file.

Unicode preserves the original ASCII character assignments for its first 128 code points. It extends the repertoire rather than giving A a different basic number.

07 | UNICODE AND ITS ENCODINGS

Unicode defines characters and code points. An encoding such as UTF-8 specifies how to represent those code points using bytes.

Do not say that every Unicode character always uses 16 bits. UTF-8 uses one to four bytes for a code point; other Unicode encodings use different rules.

How UTF-8 represents example code points
CharacterUTF-8 bytes in hexBytes used
A411
éC3 A92
中E4 B8 AD3
😀F0 9F 98 804

The ASCII characters each use one byte in UTF-8. Other symbols may use more. A visible symbol can sometimes be made from more than one code point, so counting what appears on screen is not always the same as counting code points or bytes.

For this syllabus, the central comparison is that Unicode can represent a greater range of characters than ASCII. The encoding distinction helps keep that explanation accurate.

08 | WHEN THE READER USES THE WRONG RULES

If a file is encoded one way but decoded using incompatible rules, text can appear as unexpected characters. The binary data needs an agreed interpretation.

For example, UTF-8 encodes é with bytes C3 A9. A reader treating those bytes as Windows-1252 characters can show “é” instead.

A missing font glyph is a different problem: a correctly decoded character may appear as a box if the font or renderer cannot display it. Unicode support does not guarantee every font draws every symbol.

09 | TRY IT: INSPECT YOUR TEXT

Enter a short message. The explorer shows code points and UTF-8 bytes. Try A, a, a space, é and an emoji. Use fictional text; nothing is sent or saved.

MATCH THE TEXT TERM

Terminology

Terminology

Character

A text element such as a letter, digit, punctuation mark or space.

Character set

A repertoire of characters with agreed assigned codes.

Character code

A numeric identifier associated with a character.

ASCII

American Standard Code for Information Interchange; standard ASCII uses seven-bit codes.

Unicode

A standard supporting a broad range of writing systems and symbols.

Code point

A number assigned in Unicode, commonly written using U+ notation.

Encoding

Rules or a process representing text as data for storage or transmission.

Decoding

Interpreting encoded data back into characters.

UTF-8

A Unicode encoding using one to four bytes for a code point.

Font

A set of visual forms used to draw characters.

Control code

A code representing a control function rather than a printable symbol.

Byte

A group of eight bits.

Questions

Questions

CHECK YOUR UNDERSTANDING

Select all correct choices. Each exact set earns one point.

1. Why is text represented in binary?
2. Which statements about standard ASCII are correct?
3. Which statements are true?
4. Why is Unicode useful for an international messaging app?
5. Which are true of Unicode and UTF-8?
6. Using the supplied table, which byte represents A?
7. Which can cause text display problems?
8. Which statements about the binary groups for CAT are true?

EXPLAIN AND APPLY

Write your answers before revealing each sample. These are practice questions, not official examination mark schemes.

1. Explain how a computer represents a short text message.

2. State one reason a messaging app uses Unicode rather than only ASCII.

3. Using the Learn table, encode CAB as three binary bytes.

4. Explain the difference between the text character 0 and the integer value 0.

5. Explain why changing a font need not change the text’s character codes.

6. Correct this claim: Unicode always stores every character in 16 bits.

INVESTIGATE | A MESSAGE FOR EVERYONE

Plan a messaging app for students who use different languages. List the kinds of characters it needs, explain why ASCII alone is insufficient, and distinguish choosing Unicode from choosing a font.

Use the explorer to compare A, é, 中 and 😀. Record each code point and UTF-8 byte count. Explain why counting visible symbols is not a reliable byte-size calculation.

Flashcards

Flashcards

Click to flip. Select the ideas you need to revisit.

0 cards selected for revision.

    Selections are kept while this page is open.

    Workbook

    Workbook

    COMING SOON

    The workbook for 1.2.1 Representing Text is coming soon. Complete the text-representation activities and keep your HTML and CSS files.