What is NFC (Canonical Composition)?

Normalization Form C: kanonik olarak ayrıştırıp sonra yeniden birleştirerek en kısa formu üretir. Veri depolama ve değişimi için önerilir; web standart formudur.

What is NFKC (Compatibility Composition)?

Normalization Form KC: uyumluluk ayrıştırması ardından kanonik birleştirme. Görsel olarak benzer karakterleri birleştirir (ﬁ→fi, ²→2, Ⅳ→IV). Tanımlayıcı karşılaştırması için kullanılır.

What is NFKD (Compatibility Decomposition)?

Normalization Form KD: yeniden birleştirme olmadan uyumluluk ayrıştırması. En agresif normalleştirme, en fazla biçimlendirme bilgisini kaybeder.

What is Kanonik denklik?

Anlamsal olarak özdeş olan ve eşit kabul edilmesi gereken iki karakter dizisi. Örnek: é (U+00E9) ≡ e + ◌́ (U+0065 + U+0301).

Algoritmalar

Normalleştirme

Unicode metnini standart kanonik forma dönüştürme işlemi. Dört form: NFC (birleştirilmiş), NFD (ayrıştırılmış), NFKC (uyumluluk birleştirilmiş), NFKD (uyumluluk ayrıştırılmış).

2022-06-01 · Updated 2024-09-26

Why the Same Text Can Look Identical But Be Different

Consider the letter é. You can encode it two ways: as a single precomposed character é (U+00E9, LATIN SMALL LETTER E WITH ACUTE) or as the sequence e (U+0065) followed by the combining acute accent ´ (U+0301). Both render identically, both are valid Unicode — but they are different byte sequences and will not compare as equal with a naive string comparison.

This is the core problem Unicode Normalization solves. Without normalization, the same word typed on macOS (which prefers decomposed forms) can fail to match the same word stored on a Linux system (which may prefer composed forms). Searching, sorting, deduplication, and hashing all break when you have silent encoding differences.

The Four Normal Forms

Unicode defines four normalization forms, each serving different needs:

Form	Full Name	What it does
NFC	Canonical Decomposition + Canonical Composition	Decompose then recompose — most compact canonical form
NFD	Canonical Decomposition	Fully decompose to base + combining marks
NFKC	Compatibility Decomposition + Canonical Composition	Like NFC but also folds compatibility variants
NFKD	Compatibility Decomposition	The most aggressive decomposition

The "K" variants additionally fold compatibility characters — characters that are semantically equivalent but visually or historically distinct, such as ﬁ (fi ligature, U+FB01) → fi, or ² (superscript 2, U+00B2) → 2.

Using Normalization in Python

Python's unicodedata module provides normalization through a single function:

import unicodedata

text = "caf\u00e9"          # café with precomposed é (NFC)
nfd = unicodedata.normalize("NFD", text)
print(len(text))             # 4
print(len(nfd))              # 5 (e + combining acute)

# Roundtrip
assert unicodedata.normalize("NFC", nfd) == text

# Checking which form a string is already in
print(unicodedata.is_normalized("NFC", text))   # True
print(unicodedata.is_normalized("NFD", text))   # False

A safe comparison pattern for user-facing text:

def normalize_for_comparison(s: str) -> str:
    return unicodedata.normalize("NFC", s.casefold())

Quick Facts

Property	Value
Unicode standard	The Unicode Standard, Section 3.11
Python module	`unicodedata.normalize(form, string)`
Valid form names	`"NFC"`, `"NFD"`, `"NFKC"`, `"NFKD"`
Web standard	W3C recommends NFC for all web content
macOS file system	HFS+ stores filenames in NFD
Idempotency	Applying normalization twice gives the same result
Related concept	Canonical equivalence, compatibility equivalence

İlgili Terimler

NFC (Canonical Composition) NFD (Canonical Decomposition) NFKC (Compatibility Composition) NFKD (Compatibility Decomposition) Kanonik denklik

Algoritmalar içinde daha fazlası

Bileşim dışlama

Başlatıcı olmayan ayrıştırmayı önlemek ve algoritmik kararlılığı sağlamak için kanonik birleştirmeden (NFC) …

Case Folding

Mapping characters to a common case form for case-insensitive comparison. More comprehensive …

Cümle sınırı

Unicode kurallarına göre cümleler arasındaki konum. Noktalara göre bölmekten daha karmaşıktır — …

Grapheme Cluster Boundary

Rules (UAX#29) for determining where one user-perceived character ends and another begins. …

Harmanlama algoritması

Unicode dizilerini çok seviyeli karşılaştırma kullanarak karşılaştırma ve sıralama için standart algoritma: …

Kelime sınırı

Unicode kelime kesme kurallarına göre belirlenen kelimeler arasındaki konum. Boşluklara göre basit …

Metin bölümleme

Metinde sınır bulma algoritmaları: grafem kümesi, kelime ve cümle sınırları. İmleç hareketi, …

NFC (Canonical Composition)

Normalization Form C: kanonik olarak ayrıştırıp sonra yeniden birleştirerek en kısa formu …

NFD (Canonical Decomposition)

Normalization Form D: yeniden birleştirmeden tamamen ayrıştırır. macOS HFS+ dosya sistemi tarafından …

NFKC (Compatibility Composition)

Normalization Form KC: uyumluluk ayrıştırması ardından kanonik birleştirme. Görsel olarak benzer karakterleri …

← Sözlüğe Geri Dön