Research · Open data · OpenNHE Technologies

Emotion labeling research and an open dataset

Pratham Prateek Mohanty · Masters' Union · OpenNHE Technologies

Three technical papers on small LoRA adapters that label emotions in text, and the public synthetic dataset behind the Elysium X 150 FR paper. Every limit is stated next to its number.

3technical papers with DOIs
4,000labeled target turns in the open dataset
150emotion and appraisal coordinates
CC BY 4.0dataset license

Papers

Open: model, data, code

Elysium X 150 FR

A small LoRA adapter (73.9 MB on Qwen2.5-1.5B-Instruct) for sparse, per-speaker emotion and appraisal labeling over a 150-coordinate schema. On a frozen 78-row synthetic test it reaches micro-F1 0.7251 and 49 of 78 exact label sets, agreement with a reviewed synthetic teacher only.

No external benchmark, 117 of 150 coordinates seen in training, English only. No state-of-the-art claim.

Open: adapter weights on Hugging Face

Elysium X 20 FR

A small LoRA adapter (18.5M trainable parameters) on Qwen2.5-1.5B-Instruct that reads one English message and returns structured JSON: GoEmotions labels with rater-agreement intensities, plus valence and arousal. On 1,500 GoEmotions test examples it reaches micro-F1 0.570 against 0.304 for always predicting neutral.

One seed, one run, no confidence intervals, English Reddit text only. First version, no state-of-the-art claim.

Paper only: model and data are proprietary

Elysium X 500 FR

A first baseline for multilingual emotion labeling over a 500-coordinate schema in 12 languages, measured when the emotion word is hidden. The headline is the hard one: 19.4% exact match on 403 no-name rows. With the word present it is 95.1%, which mostly reflects recognising a label word.

Template-generated synthetic data, possible translation leakage, one seed. The weights, data and code are not public.

Open data: Elysium X 150 FR Emotion Dataset

4,000 labeled target turns from 1,000 four-turn English dialogues, annotated on the 150-coordinate schema. The dialogues are synthetic, written by an AI from scenario prompts. They are not real people's conversations.

  • Labels came from the author's ChatGPT labeling workflow and were personally reviewed by him. That is owner-reported review, not independent multi-annotator gold.
  • 1,401 of 4,000 rows have no label, meaning no expressed coordinate for that speaker at that turn.
  • English only. Not a clinical, diagnostic or safety dataset.
  • The model's 78 test rows are marked test in model_split. Do not train on them if you want to compare with the paper.

Explore the data

A random sample of 250 rows from the open dataset (210 labeled, 40 unlabeled), shown with the label names. The target turn is outlined. Full data is in the dataset repo.

The 150-coordinate schema

Ten families of fifteen coordinates. Each has an operational definition and a note on what is not enough on its own, in taxonomy.json in the dataset repo.

Cite

Mohanty, P. P. (2026). Elysium X 150 FR: A Small LoRA Adapter for Sparse, Per-Speaker Emotion and Appraisal
Labeling over a 150-Coordinate Schema. Zenodo. https://doi.org/10.5281/zenodo.23155240

Mohanty, P. P. (2026). Elysium X 500 FR: A First Public Baseline for Multilingual Emotion Labeling over a
500-Coordinate Schema When the Emotion Word Is Hidden. Zenodo. https://doi.org/10.5281/zenodo.23159402

Mohanty, P. P. (2026). Elysium X 20 FR: A Small LoRA Adapter That Turns One Message into a Structured Emotion
Appraisal, Trained on GoEmotions. Zenodo. https://doi.org/10.5281/zenodo.23161139

Mohanty, P. P. (2026). Elysium X 150 FR Emotion Dataset (CC BY 4.0).
https://huggingface.co/datasets/open-nhe/Elysium-X-150-FR-dataset