Skip to content

Hebrew ML Datasets Navigator

Trusted87/100
Before deciding whether to install, talk to the skill

Navigate the fragmented landscape of Hebrew and Yiddish ML datasets and models. Covers ivrit.ai (20K+ hours of Hebrew audio, whisper-large-v3 ASR variants, Yiddish models), Dicta (DictaLM 3.0 LLM family, DictaBERT variants, HeQ reading comprehension), the Israeli National NLP Program / NNLP-IL (HebrewSentiment, HebNLI), AlephBERT, and Knesset Plenums. Helps researchers and ML engineers pick the right dataset for a task by use case, license (commercial vs research), Hebrew register coverage, and model-dataset pairing. Use when choosing training data for a Hebrew NLP or ASR project, verifying license compatibility for a commercial product, finding a baseline model for a Hebrew downstream task, or exploring Yiddish ML resources. Do NOT use for Arabic NLP, general HuggingFace dataset discovery, or Hebrew OCR dataset selection (use hebrew-ocr-forms).

Trust score 87/100 (Trusted) · 34+ installs · 4 GitHub contributors · MIT license

The Problem

The Israeli ML community punches above its weight, but the datasets and models are scattered. ivrit.ai publishes world-class Hebrew speech corpora on one HuggingFace org, Dicta publishes Hebrew LLMs and BERT variants on another, the Israeli National NLP Program maintains benchmarks under HebArabNlpProject. Licenses vary from fully commercial-friendly to research-only. A researcher trying to pick the right combination for fine-tuning a Hebrew sentiment classifier on customer support chat for a commercial product has to hunt across five orgs and read every dataset card.

skills-ilskills-ilDeveloper Tools
1.0.5MITGitHub
34installs1,699views
0Write a Review

How to use this skill

Not sure how? Read the guide
  1. 1. Click "Download ZIP" to download the skill files.
  2. 2. Open Claude Desktop and go to Customize > Skills.
  3. 3. Click "+" and select "Upload a skill", then upload the ZIP file.
  4. 4. Start a new conversation. The skill will activate automatically when relevant.
A new version released? How to update your installed skill
Developers? Install via command line (CLI)
npx skills-il add skills-il/developer-tools@v1.0.5-hebrew-ml-datasets-navigator --skill hebrew-ml-datasets-navigator -a claude-code

When to Apply

  • When choosing training data for a Hebrew NLP or ASR project
  • When verifying license compatibility for commercial use of a dataset
  • When looking for a baseline model for a specific Hebrew task
  • When building a Hebrew transcription stack and need to know what ivrit.ai offers
  • When researching or building something in Yiddish and need to find resources

Try These Prompts

Commercial sentiment

I want to train a sentiment classifier on Hebrew customer support chat for a commercial SaaS product. Which dataset should I use, which starting model, and what does the license say about attribution?

Hebrew podcast transcription

I am building a Hebrew podcast transcription product. What does ivrit.ai offer, which ASR model should I use in production with low latency, and how do I handle multiple speakers?

Small Hebrew LLM

I need a Hebrew LLM that runs on consumer hardware (16GB VRAM max) for a Hebrew product. What does Dicta offer, what are the size differences, and what are the upstream licenses?

Yiddish ML

I am researching Yiddish and looking for datasets and models for speech recognition and text processing. What is available in 2026 and what are the licenses?

Frequently Asked Questions

Changelog

v1.0.5

Update: OSCAR-2301 access is temporarily suspended (users routed to CulturaX/FineWeb-2), and the ivrit.ai corpus size was corrected from 22K to 20K hours per the source dataset page.

Jul 6, 2026

v1.0.4

Added Jamba 1.6 and Jamba-Reasoning-3B to the model catalog, acknowledged the DictaLM 3.0 benchmark suite (Translation, Summarization, Winograd, Israeli Trivia, Diacritization), and redirected HebrewSentiment license claims to the live dataset card.

May 19, 2026

v1.0.3

Added HEBREW-MMLU, CulturaX, FineWeb-2, ParaShoot, HeSum, academic resources. Stripped 27 em dashes.

Apr 25, 2026

Related Skills

skills-ilAuthor: skills-il
v2.2.1Popular

Build and configure Make.com scenarios for Israeli business processes, including Morning (formerly Green Invoice) sync, iCount accounting, Monday.com board automation, Priority ERP data exports, WhatsApp Business messaging, and payment gateways (Cardcom, Tranzila, Grow, Bit). Covers Make.com AI Agents, the Make.com MCP server for exposing scenarios as agent tools, Israel 2026 Invoice Reform, community modules for Israeli apps, Hebrew data transformations, Data Store for VAT period tracking, and Shabbat-aware scheduling. Do NOT use for n8n workflows (use n8n-hebrew-workflows) or Zapier Zaps (use zapier-israeli-integrations).

0.01302,407
Claude CodeCursorGitHub Copilot+4
skills-ilAuthor: skills-il
v2.0.0Popular

Convert between Hebrew (Jewish) calendar and Gregorian dates, look up Israeli holidays, format dual dates for Israeli documents, and calculate Israeli business days. Use when user asks about Hebrew dates, "luach ivri", Jewish calendar, Israeli holidays, "chagim", Shabbat times, or needs dual-date formatting for Israeli forms. Do NOT use for Islamic Hijri calendar or non-Israeli holiday calendars.

0.0971,959
Claude CodeCursorGitHub Copilot+6
skills-ilAuthor: skills-il
v1.2.0Popular

Benchmark and compare LLMs on Hebrew reasoning, comprehension, sentiment, translation, and Israeli cultural knowledge. Wraps the HuggingFace Open Hebrew LLM Leaderboard tasks (HeQ, HebrewSentiment, Hebrew Winograd, translation) plus DictaLM 3.0 benchmark tasks (Summarization, Nikud, Israeli Trivia) into a reproducible evaluation harness. Runs evals against Claude, GPT, Gemini, AI21 Jamba, DictaLM, Llama, and local HuggingFace models. Produces comparison scorecards in JSON and markdown. Use when choosing an LLM for a Hebrew product, answering procurement questions about Hebrew performance, validating a fine-tuned Hebrew model, or tracking Hebrew regressions after a model upgrade. Do NOT use for Arabic NLP, ASR benchmarking, or general English benchmarks.

0.0602,968
Claude CodeCursorCodex
Found an issue with this skill?

Use at your own risk. Terms of Use · Security

Want to build your own skill? Try the Skill Creator · Submit a Skill

Reviews (0)

No reviews yet. Be the first to write one!