| single |
# Internationalized Domain Names in Applications (IDNA)
Support for [Internationalized Domain Names in Applications
(IDNA)] and [Unicode IDNA
Compatibility Processing]. It
supersedes the standard library's `encodings.idna`, which only
implements the 2003 specification, offering broader script coverage and
limiting domains with known security vulnerabilities.
## Usage
Package may be installed from [PyPI] via
the typical methods (e.g. `python3 -m pip install idna`)
For typical usage, the `encode` and `decode` functions will take a
domain name argument and perform a conversion to ASCII-compatible encoding
(known as A-labels), or to Unicode strings (known as U-labels)
respectively.
```pycon
>>> import idna
>>> idna.encode('ドメイン.テスト')
b'xn--eckwd4c7c.xn--zckzah'
>>> print(idna.decode('xn--eckwd4c7c.xn--zckzah'))
ドメイン.テスト
```
Conversions can be applied at a per-label basis using the `ulabel` or
`alabel` functions for specialized use cases.
### Compatibility Mapping (UTS #46)
This library provides support for [Unicode IDNA Compatibility
Processing] which normalizes input from
different potential ways a user may input a domain prior to performing the
IDNA
conversion operations. This functionality, known as a
[mapping], is considered by the
specification to be a local user-interface issue distinct from IDNA
conversion functionality.
For example, "Königsgäßchen" is not a permissible label as capital
letters
are not allowed. UTS #46 will convert this into lower case prior to
applying
the IDNA conversion.
```pycon
>>> import idna
>>> idna.encode('Königsgäßchen')
...
idna.core.InvalidCodepoint: Codepoint U+004B at position 1 of
'Königsgäßchen' not allowed
>>> idna.encode('Königsgäßchen', uts46=True)
b'xn--knigsgchen-b4a3dun'
>>> idna.decode('xn--knigsgchen-b4a3dun')
'königsgäßchen'
```
When performing a decode operation for display purposes, `decode()`
accepts a `display=True` argument that leaves any `xn--` label that
fails to decode unchanged. This is useful for user interface display
where a domain is in use, the A-label form can be presented when it
is not a valid IDN.
## Exceptions
All errors raised during conversion derive from the `idna.IDNAError`
base class. The more specific exceptions are:
* `idna.IDNABidiError` — raised when a label contains an illegal
combination of left-to-right and right-to-left characters.
* `idna.InvalidCodepoint` — raised when a label contains a codepoint
that is INVALID for IDNA.
* `idna.InvalidCodepointContext` — raised when a CONTEXTO or CONTEXTJ
codepoint appears in a position whose contextual requirements are
not satisfied.
Exceptions carry machine-readable attributes so that applications
do not need to parse the message: `code` is a short, stable identifier
for the rule that failed (listed below); and, when the failure can be
attributed to a particular character, `text` (the label, or domain for
UTS #46 processing, being validated), `codepoint` (the offending
codepoint as an integer) and `position` (its 1-based index within
`text`, as quoted in the message) are set. Each is `None` when it does
not apply. Message wording is not part of the API and may change.
```pycon
>>> try:
... idna.encode('Königsgäßchen')
... except idna.IDNAError as err:
... print(err.code, err.codepoint, err.position, err.text)
disallowed_codepoint 75 1 Königsgäßchen
```
| `code` | Meaning |
|---|---|
| `input_too_long` | Input exceeds the library's defensive length limit and
was not processed |
| `label_too_long` | A label exceeds 63 octets |
|