Universal Tripiṭaka: Taishō Canon Morse Transliteration Layer
Item Preview
Share or Embed This Item
data
Universal Tripiṭaka: Taishō Canon Morse Transliteration Layer
- Publication date
- 2026-02-14
- Usage
- Attribution-NonCommercial-NoDerivs 4.0 International
- Topics
- Buddhism, Tripiṭaka, Taishō Shinshū Daizōkyō, CBETA, SAT, Transliteration, Digital Humanities, Chinese Buddhist Canon, Machine-Readable Corpus, Morse
- Language
- Morse
- Rights
- Copyright Disclaimer and Fair Use Notice: Educational Purpose: These files are uploaded strictly for non-commercial, educational, and archival purposes under the Fair Use doctrine. Public Domain and Orphan Works: Many of the historical editions included here are considered public domain due to their age. Right to Removal: We hold the deepest respect for intellectual property. If you are the legal copyright holder and object to its availability, please contact us. No Commercial Intent: No profit is derived from this collection. May this Dhamma-Dana lead to the peace and wisdom of all beings.
- Item Size
- 93.8G
KIT-2920: Universal Tripiṭaka Taishō Morse Reading Layer is a transformative digital humanities infrastructure designed to bridge the gap between complex academic archives and modern accessibility. This project systematically processes the vast Chinese Buddhist Canon, primarily the Taishō Shinshū Daizōkyō based on CBETA and SAT (Saṃgha Academic Taishō) databases擁nto a streamlined, machine-friendly, and human-readable Morse transliteration layer. By stripping away the dense nesting of TEI/XML editorial markups, KIT-2920 provides flat UTF-8 text outputs that preserve the integrity of the original 85 volumes while offering a clean Reading-and-Chanting-First experience. The corpus covered the full 2920-item Taishō collection, utilizing a multi-layered dictionary architecture, including Buddhist-specific phrases and Phrase overrides葉o ensure high-precision script with tone marks. This initiative does not replace the critical editions of CBETA or SAT; rather, it builds a script bridge for developers, AI training systems, and global readers who require immediate linguistic access to canonical texts without the barriers of technical parsing. Morse:是一個致力於將傳統學術佛典轉化為現代數位化友善格式的人文計算基礎設施。本專案由 kit119 開發,針對 CBETA 與 SAT(大正藏圖像數據庫)所提供的龐大漢字文本進行系統化處理,生成具備聲 調的完整拼音層與純文本層。KIT-2920 的核心理念在於「閱讀優先」,透過移除複雜的 TEI/XML 嵌套標籤,解決了開發者與一般讀者在處理佛典時的技術門檻。專案內容涵蓋大正藏全部 85 卷(共 2920 部經),並特別保留了 SAT 與 CBETA 之間的編輯差異,不進行強制合併,確保學術誠實。系統採用多層級詞庫架構(包含佛教專用詞彙與越漢對照優化),確保轉寫過程無空白遺漏且精準。此專案並非旨在取代 CBETA 或 SAT 的學術地位,而是為開發者、人工智慧模型及全球漢語學習者建立一座跨越文字障礙的橋樑,讓傳承千年的大藏經能以更開放、更易讀的形態進入數位時代。
- Addeddate
- 2026-02-12 00:49:43
- Collection_added
-
folkscanomy_religion
folkscanomy - Identifier
- universal-tripitaka-taisho-morse
- Ocr
- tesseract 5.3.0-6-g76ae
- Ocr_autonomous
- true
- Ocr_detected_lang
- af
- Ocr_detected_lang_conf
- 1.0000
- Ocr_detected_script
- Devanagari
- Ocr_detected_script_conf
- 0.9968
- Ocr_invalid_language
- Morse
- Ocr_module_version
- 0.0.21
- Ocr_parameters
- -l bre+deu+spa+eng+afr+lat
- Page_number_confidence
- 0
- Page_number_module_version
- 1.0.5
- Ppi
- 300
comment
Reviews
5 Views
DOWNLOAD OPTIONS
IN COLLECTIONS
KIT: Knowledge for International TripitakaUploaded by kit119 on