Data sources
Origa builds on open data and models. This page lists what the app uses and the terms each derived work is distributed under.
Dictionary and readings
Dictionary entries, translations, and furigana come from JMdict / EDRDG under CC BY-SA 4.0. The dictionary project is maintained by the Electronic Dictionary Research and Development Group.
Kanji animations
Stroke order data and animations come from KanjiVG under CC BY-SA 3.0.
Tokenization
Text segmentation uses SudachiDict under the Apache-2.0 license.
OCR
Image text recognition uses NDLOCR-Lite from the National Diet Library of Japan under CC BY 4.0.
Speech recognition
Audio transcription uses Whisper under the MIT license.
Phrase audio
Listening practice recordings come from native-speaker corpora, including NHK material.
Word sets
Pre-built word sets for import come from Irodori by the Japan Foundation.
Fonts
- Cormorant Garamond — SIL Open Font License
- IBM Plex Mono — SIL Open Font License
- Noto Sans JP — SIL Open Font License
