What is Normalleştirme?

Unicode metnini standart kanonik forma dönüştürme işlemi. Dört form: NFC (birleştirilmiş), NFD (ayrıştırılmış), NFKC (uyumluluk birleştirilmiş), NFKD (uyumluluk ayrıştırılmış).

What is Kanonik denklik?

Anlamsal olarak özdeş olan ve eşit kabul edilmesi gereken iki karakter dizisi. Örnek: é (U+00E9) ≡ e + ◌́ (U+0065 + U+0301).

Algoritmalar

NFC (Canonical Composition)

Normalization Form C: kanonik olarak ayrıştırıp sonra yeniden birleştirerek en kısa formu üretir. Veri depolama ve değişimi için önerilir; web standart formudur.

2022-06-14 · Updated 2024-07-11

NFC: The Web's Default Normal Form

NFC (Normalization Form C — Canonical Composition) is the most widely used Unicode normalization form. It works in two passes: first, it decomposes all characters into their canonical base + combining mark sequences (like NFD), then it recomposes them back into precomposed characters wherever the Unicode standard defines a canonical composition.

The result is the shortest canonical representation of a string. For most Latin-script text, NFC means characters like é, ü, and ñ are stored as single code points rather than two-code-point sequences. For text that is already in NFC (pure ASCII, for instance), normalization is a no-op.

Why NFC is the Recommended Default

The W3C mandates NFC for all web content in the Character Model for the World Wide Web. Most databases, APIs, and programming environments assume NFC. HTTP headers, JSON payloads, and HTML source files are all expected to use NFC.

macOS user-space applications generally use NFC (despite HFS+ using NFD internally — the OS translates at the file system boundary). Windows and Linux also default to NFC in most contexts. If you are writing text to a file, database, or API and you want maximum interoperability, NFC is the right choice.

Python Examples

import unicodedata

# e + combining acute → é (one code point)
decomposed = "e\u0301"      # NFD form: 2 code points
composed = unicodedata.normalize("NFC", decomposed)

print(repr(decomposed))     # 'e\u0301'
print(repr(composed))       # '\xe9'  (which is é, U+00E9)
print(len(decomposed))      # 2
print(len(composed))        # 1

# Normalize user input before storing
def store_text(text: str) -> str:
    return unicodedata.normalize("NFC", text)

# NFC is idempotent
s = "caf\u00e9"
assert unicodedata.normalize("NFC", s) == s
assert unicodedata.is_normalized("NFC", s)

NFC does NOT fold compatibility characters. The fi ligature ﬁ (U+FB01) remains ﬁ under NFC. For search and identifier normalization where you want ﬁ == fi, use NFKC instead.

Quick Facts

Property	Value
Full name	Normalization Form Canonical Composition
Algorithm	NFD first, then canonical composition
Typical use	Web content, databases, API responses, user input storage
W3C standard	Required for all web content (Character Model for the WWW)
Python	`unicodedata.normalize("NFC", s)`
Handles compatibility chars?	No — use NFKC for that
Idempotent?	Yes
Comparison to NFD	Usually equal or shorter (composed chars save one code point each)

İlgili Terimler

Normalleştirme NFD (Canonical Decomposition) Kanonik denklik

Algoritmalar içinde daha fazlası

Bileşim dışlama

Başlatıcı olmayan ayrıştırmayı önlemek ve algoritmik kararlılığı sağlamak için kanonik birleştirmeden (NFC) …

Case Folding

Mapping characters to a common case form for case-insensitive comparison. More comprehensive …

Cümle sınırı

Unicode kurallarına göre cümleler arasındaki konum. Noktalara göre bölmekten daha karmaşıktır — …

Grapheme Cluster Boundary

Rules (UAX#29) for determining where one user-perceived character ends and another begins. …

Harmanlama algoritması

Unicode dizilerini çok seviyeli karşılaştırma kullanarak karşılaştırma ve sıralama için standart algoritma: …

Kelime sınırı

Unicode kelime kesme kurallarına göre belirlenen kelimeler arasındaki konum. Boşluklara göre basit …

Metin bölümleme

Metinde sınır bulma algoritmaları: grafem kümesi, kelime ve cümle sınırları. İmleç hareketi, …

NFD (Canonical Decomposition)

Normalization Form D: yeniden birleştirmeden tamamen ayrıştırır. macOS HFS+ dosya sistemi tarafından …

NFKC (Compatibility Composition)

Normalization Form KC: uyumluluk ayrıştırması ardından kanonik birleştirme. Görsel olarak benzer karakterleri …

NFKD (Compatibility Decomposition)

Normalization Form KD: yeniden birleştirme olmadan uyumluluk ayrıştırması. En agresif normalleştirme, en …

← Sözlüğe Geri Dön