Her yazı sistemindeki her karaktere benzersiz bir sayı (kod noktası) atayan evrensel karakter kodlama standardı. Sürüm 16.0, 154.998 atanmış karakter içerir.

What is Kod noktası?

Unicode kod alanındaki sayısal değer (U+0000 ile U+10FFFF arası), U+XXXX şeklinde yazılır. Tüm kod noktaları karakterlere atanmış değildir.

What is Karakter olmayan?

Dahili kullanım için kalıcı olarak ayrılmış kod noktaları (toplam 66): U+FDD0–U+FDEF ve her düzlem için U+nFFFE/U+nFFFF. Metinde geçerlidir ancak harici olarak paylaşılmamalıdır.

Unicode Standardı

Özel kullanım alanı

Kuruluşların kendi karakterlerini atayabileceği ayrılmış aralıklar: BMP PUA (U+E000–U+F8FF) ve Düzlem 15 ve 16'daki Ek PUA'lar.

2021-06-21 · Updated 2024-10-08

What is the Private Use Area?

The Private Use Area (PUA) refers to three ranges of Unicode code points that are permanently reserved for applications to define their own characters. Unlike most of the Unicode code space, PUA code points will never be assigned official characters by the Unicode Consortium. Instead, any organization can use them for proprietary characters — custom icons, corporate logos, game symbols, or glyphs not yet in Unicode.

There are three PUA regions in Unicode:

Name	Range	Size
BMP Private Use Area	U+E000–U+F8FF	6,400 code points
Supplementary Private Use Area A	U+F0000–U+FFFFF	65,534 code points
Supplementary Private Use Area B	U+100000–U+10FFFF	65,534 code points

Total: 137,468 code points — by far the largest reserved region in Unicode.

How the PUA is Used

Because PUA code points have no standard meaning, their interpretation is entirely up to the parties exchanging the text. This requires both sides to agree on a mapping — typically through a custom font that maps PUA code points to specific glyphs.

Common use cases:

Icon fonts — Font Awesome, Material Icons, and similar libraries map their icons to PUA code points (e.g., U+F000+ for Font Awesome). The font renders the PUA code point as the intended icon.
Corporate logo characters — Companies sometimes use PUA slots for brand marks in specialized documents.
Pre-standardization characters — Klingon, Tengwar (Tolkien's Elvish script), and other scripts not yet in Unicode have community-defined PUA assignments (the ConScript Unicode Registry, CSUR).
Regional/historic writing systems — Script communities waiting for official Unicode approval use the PUA for interoperability within their community.

The Interoperability Problem

PUA usage is inherently non-interoperable across different applications or organizations unless both use the same font and the same mapping. A PUA code point U+E001 might be a "thumbs up" icon in one font and a currency symbol in another. When text with PUA characters is exchanged between systems using different fonts, the result is meaningless glyphs.

# PUA code points have no official name
import unicodedata

cp = 0xE001  # PUA code point
try:
    name = unicodedata.name(chr(cp))
except ValueError as e:
    print(e)  # no such name

category = unicodedata.category(chr(cp))
print(category)  # "Co" (Private Use)

PUA in Emoji History

Before emoji were standardized in Unicode 6.0 (2010), Japanese mobile carriers (DoCoMo, KDDI, SoftBank) each used their own PUA encodings for emoji. DoCoMo used the range U+E63E–U+E757; SoftBank used a different range. This is why early cross-carrier emoji were garbled — each carrier had a different PUA mapping. Unicode 6.0 unified these into standardized code points.

Detecting PUA Characters

import unicodedata

def is_pua(char: str) -> bool:
    return unicodedata.category(char) == "Co"

print(is_pua("\uE001"))     # True (BMP PUA)
print(is_pua("\U000F0001")) # True (Supplementary PUA A)
print(is_pua("A"))          # False

Common Pitfalls

Assuming PUA characters are portable: Never embed PUA characters in data exchanged with external systems without documenting the required font/mapping.

Font Awesome characters in databases: Storing Font Awesome PUA icons in a database works only if the rendering system also uses Font Awesome. On different systems, PUA values appear as blank boxes or unrelated glyphs.

Quick Facts

Property	Value
BMP PUA range	U+E000–U+F8FF
Supplementary PUA A	U+F0000–U+FFFFF
Supplementary PUA B	U+100000–U+10FFFF
Total PUA code points	137,468
General category	Co (Private Use)
Official character assignment	Never — permanently private
Common use	Icon fonts (Font Awesome, Material Icons)
Registry for scripts	CSUR (ConScript Unicode Registry)

İlgili Terimler

Unicode Kod noktası Karakter olmayan

Unicode Standardı içinde daha fazlası

Atanmamış kod noktası

Henüz hiçbir Unicode sürümünde bir karaktere atanmamış kod noktası, Cn (Atanmamış) olarak …

Atanmış karakter

Bir Unicode sürümünde karakter ataması yapılmış kod noktası. Unicode 16.0 itibariyle, 1.114.112 …

Ayrılmış kod noktası

Gelecekteki standardizasyon için ayrılmış kod noktası; kalıcı olarak ayrılan noncharacter'lardan ve kullanıcı …

Basic Multilingual Plane (BMP)

Düzlem 0 (U+0000–U+FFFF), Latin, Yunan, Kiril, CJK, Arap ve çoğu sembol dahil …

CJK

Çince, Japonca ve Korece — Unicode'da birleştirilmiş Han ideograf bloğu ve ilgili …

Düzlem

65.536 kod noktasından oluşan bitişik blok. Unicode'da 17 düzlem vardır (0–16): Düzlem …

Ek düzlem

Düzlem 1–16 (U+10000–U+10FFFF), emoji, tarihi yazılar, CJK uzantıları ve müzik notasyonu içerir. …

Han Unification

The process of mapping Chinese, Japanese, and Korean ideographs that share a …

Hangul Jamo

The individual consonant and vowel components (jamo) of the Korean Hangul writing …

ISO 10646 / Universal Character Set

Unicode ile senkronize edilmiş, aynı karakter repertuvarını ve kod noktalarını tanımlayan ancak …

← Sözlüğe Geri Dön