Skip to content

Repository files navigation

ORCA: Open-ended Response Correctness Assessment for Audio Question Answering

Authors: Šimon Sedláček¹, Sara Barahona², Cecilia Bolaños³, Laura Herrera-Alarcón², Sathvik Udupa¹, Fernando López², Allison Ferner⁴, Alicia Lozano-Diez², Bolaji Yusuf¹, Santosh Kesiraju¹, Ramani Duraiswami⁵, Jan Černocký¹

¹Speech@FIT, Brno University of Technology, Czechia  ·  ²Universidad Autónoma de Madrid, Spain  ·  ³University of Buenos Aires, Argentina  ·  ⁴Tufts University, USA  ·  ⁵University of Maryland, USA

Paper: https://doi.org/10.1162/TACL.a.798 -- accepted to Transactions of the ACL 2026

ORCA example

Dataset: BUT-FIT/orca-audio-qa-annotations

Corrected benchmarks:

The corrected and rephrased MMAU and MMAR benchmarks can be used for open-ended evaluation.

Models:


Abstract: Reliable assessment of the abilities of large audio language models (LALMs) is essential to advancing the state of the art. As benchmarks rapidly evolve to incorporate complex reasoning and subjective tasks, they increasingly necessitate open-ended responses from LALMs. We present Open-ended Response Correctness Assessment (ORCA) — a reliable and lightweight model-based approach for answer correctness and disagreement modeling. We employ a three-stage annotation pipeline combining human judgment, structured feedback, and human-AI correction, yielding 9,663 annotations across 3,699 question-answer pairs from 15 LALMs on three audio understanding and reasoning benchmarks (achieving a Krippendorff's alpha of 0.82). Our experiments employing curriculum learning show that ORCA models achieve a Spearman correlation of 0.91 with average human correctness ratings on seen benchmarks and generalize to unseen benchmarks with a score of 0.85, outperforming several LLM judge baselines including Gemini 2.5 Flash. Furthermore, we demonstrate that ORCA's predicted variance correlates strongly with human disagreement, allowing it to effectively identify problematic benchmark items.


Installation

# Python 3.12+ required
pip install -e .

Using uv (recommended for development):

uv venv --python 3.12 && source .venv/bin/activate
uv pip install -e ".[dev]"

Inference with pre-trained models

The download_and_infer.py script handles everything — model download, data download, inference, and score clamping:

python download_and_infer.py --model gemma-4b --benchmark mmau-pro

Available options:

  • --model: gemma-4b (default: olmo-1b), llama-3b, olmo-1b
  • --benchmark: mmau-pro (default), mmau-mmar
  • --stages 1 2 3 4: run all steps (default) or a subset — e.g. --stages 3 4 to skip downloading
  • --var_threshold: override the clamping threshold (default: per-model calibrated value)

Results are written to results/<model>/<benchmark>/final_result.jsonl (rating_orca, variance_orca per item) and score_clamp.yaml (Spearman, Kendall, MAE).

Manual steps

# 1. Download a model
pip install huggingface_hub
hf download BUT-FIT/orca-gemma-3-4b-it-multinomial --local-dir orca-gemma-4b

# 2. Prepare input JSONL (one item per line)
# Required fields: question, reference, candidate, rationale
# Optional: ratings (list of int 1–5) to get Spearman/Kendall/MAE metrics

# 3. Run inference
orca-infer --model_path orca-gemma-4b/model --data_jsonl your_data.jsonl --output_dir results/

# 4. Apply clamping
python -m orca_score.clamp --apply_dir results/your_data --var_threshold 0.05

Training

orca-train \
    --train_data /path/to/train.jsonl \
    --val_data   /path/to/dev.jsonl \
    --model allenai/OLMo-2-0425-1B-Instruct \
    --score_type multinomial \
    --lora_rank 128 \
    --output_dir output/

Training data is available at BUT-FIT/orca-audio-qa-annotations. To reproduce the full three-stage curriculum, train sequentially on stage1_pretrainstage2_benchmarkstage3_mmau_mmar, passing --load_checkpoint between stages.

Run orca-train --help and orca-infer --help for the full argument reference.

Model architecture

ORCA fine-tunes a pre-trained LM (Gemma, Llama, OLMo, …) with a LoRA adapter and a small linear scoring head that models the distribution over a 5-point Likert scale, from which a continuous correctness score in [0, 1] and a variance estimate are derived. Score clamping (orca_score/clamp.py) optionally hard-clamps extreme, low-variance predictions to 0 or 1.

Repository structure

orca_score/
├── model.py    # ORCA model
├── train.py    # Training loop (orca-train)
├── infer.py    # Inference (orca-infer)
├── data.py     # Data loading
├── clamp.py    # Score clamping
├── cli.py      # Entry points
└── utils.py    # Helpers
download_and_infer.py   # End-to-end convenience script

Citation

@article{sedlacek-etal-2026-orca,
    author = {Sedláček, Šimon and Barahona, Sara and Bolaños, Cecilia and Herrera-Alarcón, Laura and Udupa, Sathvik and López, Fernando and Ferner, Allison and Yusuf, Bolaji and Lozano-Diez, Alicia and Kesiraju, Santosh and Duraiswami, Ramani and Černocký, Jan},
    title = "{ORCA: Open-ended Response Correctness Assessment for Audio Question Answering}",
    journal = {Transactions of the Association for Computational Linguistics},
    volume = {14},
    pages = {2213-2233},
    year = {2026},
    month = {08},
    abstract = {Reliable assessment of the abilities of large audio language models (LALMs) is essential to advancing the state of the art. As benchmarks rapidly evolve to incorporate complex reasoning and subjective tasks, they increasingly necessitate open-ended responses from LALMs. We present Open-ended Response Correctness Assessment (ORCA)—a reliable and lightweight model-based approach for answer correctness and disagreement modeling. We employ a three-stage annotation pipeline combining human judgment, structured feedback, and human-AI correction, yielding 9,663 annotations across 3,699 question-answer pairs from 15 LALMs on three audio understanding and reasoning benchmarks (achieving a Krippendorff’s alpha of 0.82). Our experiments employing curriculum learning show that ORCA models achieve a Spearman correlation of 0.91 with average human correctness ratings on seen benchmarks and generalize to unseen benchmarks with a score of 0.85, outperforming several LLM judge baselines including Gemini 2.5 Flash. Furthermore, we demonstrate that ORCA’s predicted variance correlates strongly with human disagreement, allowing it to effectively identify problematic benchmark items.},
    issn = {2307-387X},
    doi = {10.1162/TACL.a.798},
    url = {https://doi.org/10.1162/TACL.a.798},
    eprint = {https://direct.mit.edu/tacl/article-pdf/doi/10.1162/TACL.a.798/2624638/tacl.a.798.pdf},
}

License

MIT License — see LICENSE.

Contact

About

Open-ended Response Correctness Assessment for Audio Question Answering

Resources

Stars

6 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages