Part of LinguistPro
Reading Room
A bilingual library — the Hebrew canon with morphology on tap.
A whole library, not a handful of passages
Where the LinguistPro editor is "bring any text you like," the Reading Room is "read the canon." It turns the entire public-domain Hebrew corpus into a bilingual, guided reading library where every word is analyzable — built on the open scholarship of Project Ben-Yehuda (the Hebrew-literature counterpart to Project Gutenberg), whose open corpus we catalogue, translate and enrich. It ships inside LinguistPro and shares its privacy-first engine, but it's its own learning surface, with its own way in.
The idea
26,455 works are catalogued, organized the way scholars think about Hebrew literature: a Period → Author → Work drill (in parity with benyehuda.org). You can move from the Biblical layer through the Medieval period (the Golden Age of Sepharad), the Haskalah, the Tehiya (the revival of Hebrew as a living language), the British Mandate era, and the Modern canon — across poetry, prose, articles, memoir, letters, fables and drama.
The point isn't just access — it's a graded path. A "Where to start" shelf surfaces short texts built from high-frequency vocabulary, an easy way in; a personalized "next for you" engine recommends the next text just beyond your current level (the i+1 principle from second-language acquisition), computed entirely on-device from your own vocabulary. So a learner can begin with a few-line poem and end, text by text, reading the canon in the original — with translation, niqqud and full morphology always one tap away.
Honest, served on open
796 works are fully ready to read today: machine-translated with Gemini and vocalized with Dicta niqqud, with provenance on every card — each is labelled "machine translation" and "no audio yet", never silently dressed up as the hand-curated canon. The rest of the 26k corpus is honestly marked "translation later" and moves to "ready" over time as the translation and morphological enrichment progress.
Technically it stays true to LinguistPro's privacy-first design. The library is catalogue-driven and served-on-open: a thin index is cached up front, and a work only materializes into your browser-local storage when you open it — so a 26,000-work corpus never weighs down the device, and nothing about what you read leaves it. Tap any word for offline morphological analysis; switch on a precise "context mode" (Dicta) when you want disambiguation in a specific sentence.
Built on Google Cloud
Gemini does the translation of the corpus into Russian (and beyond), and Cloud Translation v3 backs the neural translation pipeline — the heavy compute that makes a 26k-work bilingual library feasible for a solo team.
What's next: from the literary canon to the working canon
Planned / on the roadmap — not yet shipped
The Reading Room is the same machine pointed at a fixed canon — the public-domain Hebrew literature of Project Ben-Yehuda. We want to point that same pipeline at a different kind of canon: the practical, technical knowledge a newcomer needs to qualify for a skilled trade in Israel. We're calling this planned track Reading Room — Trades.
The need is concrete. Russian-speaking olim (new immigrants) training for technical and production professions face a double barrier — learning the trade and learning it in Hebrew at the same time. There are growing vocational programs aimed at this population, including practical mechanical-engineering and production tracks. Reading Room — Trades would give those learners bilingual Hebrew–Russian study aids — worked problem-books with solutions — so they can study the material in both languages at once, with the terminology, niqqud, audio and morphology the LinguistPro engine already provides. We'd build these as study aids for learners, not as any program's official materials, and we claim no affiliation, endorsement or accreditation.
The first intended corpus is a Hebrew–Russian "Materials Science — problem book with worked solutions." The founder holds a mechanical-engineering degree and a master's, so the technical source texts and solutions would be personally curated and proofread by someone qualified in the field before the pipeline ever translates or vocalizes them — the same honest discipline as the rest of the Reading Room, where machine output is labelled as machine output and never dressed up as hand-checked. Further vocational and production-engineering subjects would follow, prioritized by the courses learners actually enrol in. None of these corpora exist yet; this is a roadmap direction, presented as such.
This is also where Google Cloud is the enabler, honestly. Today LinguistPro's runtime AI runs under each user's own Google Cloud key in the browser — which funds reading, but not the studio-side work of generating a new subject corpus. Building one is a one-time batch job on our own billing account: translating, explaining, vocalizing with niqqud and synthesizing audio for an entire problem-book before any learner opens it. That is exactly the heavy Gemini translation and explanation and Cloud Translation v3 neural translation we already run in production for the literary corpus — plus the niqqud-aware Cloud Text-to-Speech the LinguistPro engine already runs in production — with Vertex AI a planned next step for fine-tuned Hebrew models, not something in production today. As a solo, unfunded founder, the limit isn't the pipeline; it's the cost of frontier-model compute at the scale of many subjects. This is the workload Google for Startups credits would move onto Kolosei's own account — turning a roadmap into shipped study aids.