Yomimaru does not have its own Japanese dictionary. Almost no reading app does. The readings behind every furigana line and every word you have ever looked up come from projects volunteers have kept alive for decades, given away on one condition: that whoever builds on the work says where it came from.
So here is where it came from.
The dictionary
JMdict, © Electronic Dictionary Research and Development Group, used under CC BY-SA 4.0. The full terms are on the EDRDG licence page.
JMdict grew out of the EDICT project Jim Breen started in 1991, and the EDRDG has maintained it ever since. More than 200,000 of its entries sit behind the dictionary popup and the furigana engine. When Yomimaru tells you what a word means, this is who is actually telling you.
We show those entries as published, and add two things of our own. The first is JLPT levels, mapped onto entries so your study lists and the library can grade themselves. The second is a frequency rank, worked out from JMdict's own ke_pri and re_pri priority codes, which is why core N3 vocabulary now opens with 医療 and 一般 rather than whichever entry happened to sort first.
Both are adaptations of JMdict, so both carry the same CC BY-SA 4.0 licence the original does.
Stroke order
KanjiVG, © Ulrich Apel, used under CC BY-SA 3.0. The project lives at kanjivg.tagaini.net.
Every stroke that animates itself in the kana trainer comes from KanjiVG's path data. Someone hand-traced those, in the right order, for thousands of characters. Watching ぬ write itself correctly is entirely their doing.
The same data now covers kanji as well — all 2,215 characters Yomimaru knows about, drawn stroke by stroke in the order a Japanese schoolchild is taught to write them. Our copy is a redistribution of KanjiVG's own paths and carries the same CC BY-SA 3.0 licence.
Kanji readings
KANJIDIC2, © Electronic Dictionary Research and Development Group, used under CC BY-SA 4.0.
KANJIDIC2 is JMdict's companion, and it is where the on'yomi and kun'yomi of every kanji come from — the Chinese-derived reading and the native Japanese one, kept apart, which matters more than it sounds. 生 has eleven readings and knowing which kind each one is tells you when to expect it. We had them in a single merged list until now, with no way to tell them apart after the fact, and KANJIDIC2 is what un-merged it.
Our split copy is an adaptation of theirs, under the same CC BY-SA 4.0.
Pitch accent
Kanjium, © Toshiro Mifune and contributors, used under CC BY-SA 4.0. The project lives at github.com/mifunetoshiro/kanjium.
Japanese pitch accent is the thing no textbook has room for and every learner eventually discovers the hard way. 橋, 箸 and 端 are all hashi; what separates them is where the pitch drops. Almost no dictionary aimed at learners records it, because recording it means someone sitting down with 124,000 words and marking each one. Kanjium's contributors did that.
We resolved their table against JMdict offline, so the accent arrives attached to the word you looked up rather than as a second thing to download. Where two of their entries disagreed about a word and had no reading in common, we publish nothing rather than guess — that was 63 words out of 109,429.
The joined result is an adaptation of both Kanjium and JMdict, under the same CC BY-SA 4.0.
Example sentences
The Tanaka Corpus, collected by Yasuhito Tanaka and maintained by the Tatoeba Project, used under CC BY 2.0 FR.
A dictionary tells you what a word means. A sentence tells you what it does. The Tanaka Corpus is roughly 148,000 Japanese sentences with English translations, and — the part that makes it unusually valuable — a hand-built index saying which dictionary word each one is an example of. Volunteers built that index word by word.
We publish about 28,000 of those sentences, two per word at most. The rest were set aside: too long to read at a glance, no ending punctuation, or not actually containing the word they were filed under. Where a sentence could have belonged to two different words spelled the same way, we published nothing rather than pick one. The furigana is ours, generated from the corpus's own readings where it recorded them.
Tatoeba's licence asks for attribution and nothing else, so unlike everything above this section, our copy is not obliged to be share-alike — which is exactly why it lives in its own files and is never merged with the JMdict-derived ones.
Reading Japanese
Japanese does not put spaces between words, so before Yomimaru can do anything useful it has to work out where one word stops and the next starts. Two tokenizers do that job.
Kuromoji by Atilika, under the Apache License 2.0, runs inside your browser. That is why furigana appears without waiting on a server.
Sudachi by Works Applications, also under the Apache License 2.0, runs on ours, and handles the heavier analysis behind imported texts.
Type
The interface is set in Inter, M PLUS Rounded 1c, Noto Sans JP, Shippori Mincho, Yuji Syuku and Oranienbaum, all released under the SIL Open Font License 1.1.
Found a mistake?
If something here is credited wrongly, or a project that belongs on this page is missing from it, write to [email protected] and we will put it right.