📰 newsreader

hackernews score 0.89 好み 0.00 en

Unicodeを脅かす亡霊

原題: A Spectre Is Haunting Unicode

unicodejis standardghost charactersjapanese encodingcharacter encodingcjk unificationjis x 0208encoding errors
原文 ↗

日本語訳

# Unicodeを彷徨う亡霊

1978年、日本の経済産業省は、後にJIS X 0208として知られることになるエンコーディングを策定した。これは現在でも、あらゆる日本語エンコーディングの重要な参照基準となっている。しかし、JIS規格がリリースされた後、人々はある奇妙なことに気づいた。追加された文字のいくつかに明らかな出典がなく、その意味や読み方も誰にも分からなかったのである。それらがどこから来たのか、確かな者は誰もいなかった。これらは「幽霊文字」として知られるようになったものである。

長い間、幽霊文字は説明のつかない、ほとんど忘れ去られた珍現象であったが、1997年にその起源を突き止めるための調査が開始された。JIS規格のすべての文字には出典の記録があるはずであったが、たとえ存在していたとしても、通常は出典となった文書が記載されているだけで、それほど詳細なものではなかった。

出典が記載されていれば、文字の起源を辿るのは簡単だと思うかもしれないが、「出典」が何を指すのかを明確にしておく必要がある。幽霊文字の最も一般的な出典の一つは、日本の地名を網羅した『国土行政区画総覧』であった。筆者が最初そう思ったように、これはせいぜい数百ページ程度の大きな地図帳のようなものだと想像するかもしれない。しかし、最新版は全7巻に及び、各巻は約900ページもあることが判明した。ページ番号の参照なしに、たった一つの文字を追い求めることを想像してみてほしい。

困難ではあったが、幽霊文字の調査は(大部分において)その起源を突き止めることに成功した。規格策定に関わったカタログ作成者への聞き取り調査を通じて、一部の文字はカタログ作成プロセスにおける誤りとして、不注意に作り出されてしまったものであることが判明した。例えば、「妛」は「女の上に山」を記録しようとした際に生じた誤りであった。「女の上に山」という字は特定の地名に含まれており、JIS規格への採用には適していた。しかし、当時はまだそれを一文字として印刷できなかったため、「山」と「女」を別々に印刷して切り抜き、紙に貼り合わせてから複写したのである。その複写を読み取る際、二つの小さな紙の継ぎ目が一画のように見え、誤って文字に加えられてしまった。元の文字(𡚴)がJISやUnicodeに追加されたのはずっと後のことであり、私が見る限り、ほとんどのサイトでは表示されない。

結局、明確な出典も歴史的な前例も持たない文字は、たった一つ「彁」だけだった。最も可能性の高い説明は「彊」の誤読によるものだというが、具体的な事例は特定されなかった。

JIS規格が広く採用されるにつれ、これらの文字はすべてUnicodeへと入り込んだ。そしてUnicodeには、CJK統合の過程で導入された独自の幽霊文字のセットが存在している。

まとめると、1978年に一連の小さなミスが、何もないところからいくつかの文字を生み出したのである。その誤りは、定着してしまうのに十分なほど長い間、発見されずに放置された。そして今、これらの亡霊は、文字コード表の暗い隅に潜むものとして、少なくとも潜在的には、地球上のあらゆるコンピュータの一部となっている。

このペースで行けば、おそらく彼らは永遠に人類と共にあり続けるだろう。Ψ

参考文献 / 関連リンク:

- 幽霊文字 ‐ 通信用語の基礎知識 - 1997年の調査による引用を含む、最も詳細なオンラインソース。

- 大正十二年の幽霊文字 - ことばマガジン:朝日新聞デジタル - 「彊」の印刷の薄れにより、デジタル化された大正時代の新聞で「彁」が誤用された例。

- ニコニコ大百科では、それぞれを妖怪の名前として扱っている。

- 『天書(A Book from the Sky)』 - 徐氷(Xu Bing)による、造語の漢字のみを用いた手書きの本。

原文(英語)を表示

A Spectre is Haunting Unicode

In 1978 Japan's Ministry of Economy, Trade and Industry established the encoding that would later be known as JIS X 0208, which still serves as an important reference for all Japanese encodings. However, after the JIS standard was released people noticed something strange - several of the added characters had no obvious sources, and nobody could tell what they meant or how they should be pronounced. Nobody was sure where they came from. These are what came to be known as the ghost characters (幽霊文字).

For a long time the ghost characters remained an unexplained and mostly forgotten curiosity, but in 1997 an investigation was launched to discover where they had come from. While all characters in the JIS standard were supposed to have a record of their sources, even when it existed it wasn't very specific, typically just listing the document it was sourced from.

You'd think that listing the source would make tracking down the origins of the characters easy, but it's important to clarify what counts as a "source" - one of the more common sources for the ghost characters was the "Overview of National Administrative Districts" (国土行政区画総覧), a comprehensive list of place names in Japan. You might, as I initially did, imagine this to be a kind of atlas, an oversize book with at most a few hundred pages. It turns out the latest edition is a seven volume set with each volume having roughly nine hundred pages. Imagine tracking down a single character without a page reference.

Despite the difficulty, the investigation into the ghost characters was successful in discovering their origins - mostly. By interviewing the catalogers involved in the creation of the standard, the investigators established that some characters were inadvertently invented as mistakes in the cataloging process. For example, 妛 was an error introduced while trying to record "山 over 女". "山 over 女" occurs in the name of a particular place and was thus suitable for inclusion in the JIS standard, but because they couldn't print it as one character yet, 山 and 女 were printed separately, cut out, and pasted onto a sheet of paper, and then copied. When reading the copy, the line where the two little pieces of paper met looked like a stroke and was added to the character by mistake. The original character (𡚴) was not added to JIS or Unicode until much later and doesn't display on most sites for me.

In the end only one character had neither a clear source nor any historical precedent: 彁. The most likely explanation is that it was created as a misreading of the 彊 character, but no specific incident was uncovered.

Following the general adoption of the JIS standards these characters all made their way into Unicode, which has its own separate set of ghost characters introduced during CJK unification.

To sum up - in 1978 a series of small mistakes created some characters out of nothing. The errors went undiscovered just long enough to be set in stone, and now these ghosts are, at least in potential, a part of every computer on the planet, lurking in the dark corners of character tables.

At this rate they'll presumably be with humanity forever. Ψ

References / related links:

- 幽霊文字 ‐ 通信用語の基礎知識 - the most thorough online source, with citations from the 1997 investigation.

- 大正十二年の幽霊文字 - ことばマガジン:朝日新聞デジタル - an example of 彁 mistakenly used in a digitized Taisho newspaper due to a faded printing of 彊.

- Nico Nico Douga's Wiki treats each of them as the name of a youkai.

- 天书 or A Book from the Sky, a hand-printed book by Xu Bing using only made-up Chinese characters.

← 一覧に戻る