Track 2 material deadline: 24 Sep 2026, 11:59 PM IST

IndoML Launched This Independence Day! — 15th August 2026 Vaani

๐Ÿš€ Competition Started
15th August โ€” 80th Independence Day
๐Ÿ Phase 1 โ€” Half-Marathon
--d --h --m --s
๐Ÿ† Phase 2 โ€” Final Deadline
--d --h --m --s
๐Ÿ‡ฎ๐Ÿ‡ณ Independence Day Challenge • IndoML 2026

Datathon@IndoML 2026

Celebrating India's Linguistic Diversity Through ML

Noise event detection and removal in real-world Indic speech โ€” be part of India's ML community. Build something meaningful for Bharat. ๐Ÿ‡ฎ๐Ÿ‡ณ

 Open to everyone based in India โ€” students, researchers & industry professionals

โ‚น2,00,000 Prize Pool
Travel Grants
Present at IndoML

Phase 1 (Half-Marathon) Results โ€” Track 1

Top 8 โ€” Track 1 Detection Phase 1 closed 22nd September 2026

Congratulations to the Top 8 teams on the Track 1 (Noise Event Detection) Phase 1 leaderboard. The leaderboard is NOT reset โ€” all teams continue into Phase 2.

Rank Team Members (Affiliation)
๐Ÿฅ‡ 1CodeAMUShah Ahmed Shakir Abu Asim Khan (AMU), Sajid Javid (IIIT Delhi), Osama Zaheer (Binary Semantics), Ahraz Shamim (Vecmocon Technologies)
๐Ÿฅˆ 2ARCLY INDIAUmar Khan (Robonito), Mohammad Ashraf (TeleCRM), Md Mizan Ahmad (BEL), Zuhair Arif (AMU)
๐Ÿฅ‰ 3SyehsyohSaifuddin ST
4HydraManish Joshi (Independent Researcher), Neeraj Singh Aithani (Independent Researcher)
5TensorForgeMohd Huzaifa (Jamia Millia Islamia), Eraf Ali (IIT Madras), Mohd Rashid (Sofyrus Technologies), Abdul Haseeb (TCS), Mohd Azhan Kamil (Cognizant)
6Sonic BoomRohit Singhee (Walmart), Rounak Neogy (Walmart), Kalash Bhattad (Walmart), Ayush Malik (Walmart)
7BumbleBeeDipan Mandal (Research Assistant, IIT Delhi)
8liyingfaYingfa Li (Capital Medical University), Pritam Bhakat (Jawaharlal Nehru University), Shuang Liang (Capital Medical University), Yu Gu (Capital Medical University)
Track 2 (Removal) Phase 1 results will be announced separately, once the submitted models and inference code have been independently verified by the organizers.
Important โ€” Track 1 Participants Phase 2 Test Data & Leaderboard

Track 1 test data for Phase 2 is the SAME as Phase 1 โ€” that is why the leaderboard is NOT reset

For Track 1 (Noise Event Detection), the Phase 2 test set is identical to the Phase 1 test set. Nothing has changed โ€” you do not need to download any new data, and your existing pipeline will continue to work as-is. Because the data is unchanged, there is no refresh/reset of the Track 1 leaderboard โ€” all Phase 1 scores carry over and remain directly comparable with Phase 2 submissions.

Note: The changed Phase 2 test data announcement applies to Track 2 only. Track 1 participants can ignore it and simply keep improving their submissions until the Phase 2 deadline (17th October 2026, 12:00 Noon IST).
Important โ€” Track 2 Participants Model & Inference Code Submission

Track 2 teams must submit their model and inference code with the final Phase 1 submission

Track 2 Material Submission Form
Deadline: 24th September 2026, 11:59 PM IST
๐Ÿšจ Phase 2 update โ€” Track 2 test data has CHANGED. A new test set is now used for Phase 2 of Track 2. Download it on Codabench from Get Started โ†’ Files โ†’ input_data [Final Submission (Phase 2)]. If your submission is failing, please make sure you are generating and submitting your files with respect to the NEW Phase 2 test data (enhanced WAVs must match the new test clip filenames, and transcripts.jsonl must cover the new clips). Track 1 test data for Phase 2 is unchanged โ€” no leaderboard reset.

To ensure the integrity and fairness of the Track 2 evaluation โ€” and in response to concerns raised regarding submissions โ€” we are introducing an additional verification audit as a precautionary measure. Track 2 participants are requested to provide their executable model inference code and required model files so that submitted models can be independently reproduced and verified by the organizers on the evaluation data.

Along with your final Phase 1 submission, please share a ZIP package containing:

  • Model files โ€” the trained model/checkpoint and all files required to load and run the model, in the format supported by your framework (e.g., PyTorch, NeMo, etc.).
  • Requirements file โ€” requirements.txt, environment.yml, or an equivalent dependency specification.
  • Inference script โ€” a main.py script (or equivalent executable code) that runs inference. Please name the entry-point file main.py so the main file is easy to identify.
  • README โ€” clear instructions for setting up the environment and running the inference code, including the primary contact's mobile number and email.

You do NOT need to submit the SraVaani ASR model. It is simply a Hugging Face call (ARTPARK-IISc/SraVaani-1.0) โ€” you only need to provide the script that calls the Hugging Face model, and the audio is transcribed with the same ASR. This keeps the ZIP lightweight; only your enhancement model and inference code are required.

The inference script must take exactly two arguments:

  1. Input folder path โ€” containing all the test audio files.
  2. Output folder path โ€” where the inference results should be saved.

After inference, the enhanced audio files must be saved in the output folder with the same file names as the input audio files, along with the JSON transcript file in the same format as the submission format specified on Codabench.

Naming (important): Name the ZIP file and its top-level folder as your Team_name (e.g., Team_name.zip containing a Team_name/ folder). This lets us clearly identify multiple submissions and, if a reproduction issue arises, revert each team individually.

How & when to submit: Upload the ZIP package to Google Drive and submit the shareable Google Drive link through the Track 2 material submission Google Form. Deadline: 24th September 2026, 11:59 PM IST (2 days after the Phase 1 deadline, as the package needs to follow the specified format). Note: this is the deadline for the Google Form material only โ€” Phase 1 leaderboard ranks are decided by your Codabench submission at the Phase 1 deadline (22nd September, 12:00 Noon IST).
Open the Track 2 Material Submission Form
Direct link: https://docs.google.com/forms/d/e/1FAIpQLSeZ23QowJWMSS4oON-ZuW_KGgFvYFPHUVjFXY4iKYBz7M6Aig/viewform

About the Datathon

Welcome to Datathon@IndoML 2026 - a research-oriented data science competition held in conjunction with IndoML 2026. Building on the success of previous editions, this year's datathon challenges participants to tackle noise event detection and removal in real-world Indic speech - a critical problem for inclusive, robust Automatic Speech Recognition across Indian languages.

The competition is organised into two tightly coupled tracks: Track 1 (Detection) - detect noise events with precise timestamps, effectively utilising data annotated at different levels - and Track 2 (Removal) - suppress detected events while preserving the underlying speech. The dataset is a curated subset of the Vaani corpus, consisting of ~154.6 hours of real-world Indic audio with three levels of annotation quality.

Top-performing teams will be invited to attend IndoML 2026 and present their solutions to leading researchers and professionals from academia and industry. These teams will also receive exciting cash prizes.

๐Ÿ“‹ Registration is now closed. The first AMA was held on 9th September 2026 โ€” watch the recording โ†’ or view the slides โ†’

How to Participate

Step-by-Step Guide

  1. Registration (now closed): Google Form registration has closed as of 9th September 2026. Teams that have already registered can continue to make submissions on Codabench.
  2. Register on Codabench: Create an account on Codabench and join the competition for Track 1 (Detection) and/or Track 2 (Removal). All team members can join the competition, but submissions should come from one account per team.
  3. Open Track Files: For each track in Codabench, go to "Get Started" โ†’ "Files".
  4. Download Validation/Test Data: In the "Files" tab, download input_data. The validation dataset is available there for both tracks.
  5. Download Training Data: The training dataset (~154.6 hours, Gold/Silver/Bronze tiers) is hosted on Hugging Face: ARTPARK-IISc/Vaani-Noise-Event-Dataset.
  6. Build & Submit: Build your model, generate predictions, and submit on Codabench. Track 1: a predictions.jsonl file. Track 2: enhanced 16 kHz mono WAV files + transcripts.jsonl from the mandated ASR.
Join our Discord community for discussions, updates, and support throughout the competition.

Announcements

23 September 2026

Track 1 โ€” Phase 1 (Half-Marathon) Top 8 Announced

The Top 8 teams for Track 1 (Noise Event Detection), Phase 1 have been announced: CodeAMU, ARCLY INDIA, Syehsyoh, Hydra, TensorForge, Sonic Boom, BumbleBee and liyingfa. The leaderboard is not reset โ€” all teams continue into Phase 2. Track 2 results will follow after model and inference code verification. See the full list โ†’

22 September 2026

Track 1 โ€” Phase 2 Dataset Unchanged, No Leaderboard Reset

For Track 1, the Phase 2 test data is exactly the same as Phase 1. There is no new data to download, and because the dataset is unchanged the leaderboard is not refreshed/reset โ€” all Phase 1 scores carry over. The changed-test-data notice applies to Track 2 only. More details โ†’

22 September 2026

Track 2 Material Submission (Google Form) โ€” Deadline 24th September 2026, 11:59 PM IST

Track 2 teams must submit their model + inference code package via the Track 2 material submission Google Form by 24th September 2026, 11:59 PM IST. This is 2 days after the Phase 1 deadline (22nd September, 12:00 Noon IST). Phase 1 ranks are decided by your Codabench leaderboard submission at the Phase 1 deadline; the Google Form package is used for verification/reproduction. See the full requirements at the top of this page โ†’

22 September 2026

Phase 2 Live โ€” Track 2 Test Data Changed & Material Submission Form

Phase 2 has begun. For Track 2, the test data has changed for Phase 2 โ€” download the new set from Codabench at Get Started โ†’ Files โ†’ input_data [Final Submission (Phase 2)]. If your submission is failing, ensure your enhanced WAVs and transcripts.jsonl are generated with respect to the new Phase 2 test data. Track 1 test data is unchanged and the leaderboard is not reset.

Track 2 code/model submission: submit your material (inference code, model files, README) via the Track 2 material submission Google Form โ†’ See the full requirements at the top of this page โ†’

18 September 2026

Both Phase Deadlines Extended by 2 Days

Due to Codabench downtime, both deadlines have been extended by 2 days for fairness: Phase 1 (Half-Marathon) from 20th to 22nd September 2026, and Phase 2 (Final) from 15th to 17th October 2026 โ€” both at 12:00 Noon IST. All other timelines remain unchanged. Thank you for your patience!

18 September 2026

Track 2 โ€” Model & Inference Code Submission Required

To ensure the integrity and fairness of the Track 2 evaluation, all Track 2 teams must submit their executable inference code (a main.py) and required model files along with their final Phase 1 submission, so that models can be independently reproduced and verified by the organizers. Name the ZIP as your Team_name and include the primary contact's mobile & email in the README. You do not need to submit the SraVaani ASR model โ€” it's just a Hugging Face call; simply include the script that calls it (audio is transcribed with the same ASR). See the full submission requirements at the top of this page โ†’

9 September 2026

First AMA Session Held โ€” Recording & Slides Available

Our first Ask Me Anything (AMA) session was held on 9th September 2026, walking through the task, dataset, baselines and evaluation for both tracks, followed by a live Q&A. Registration is now closed.

September 2026

Validation Dataset Released!

The validation dataset is now available for both tracks in Codabench under Get Started โ†’ Files. Use it to tune and benchmark your models before the final test phase.

August 2026

Dataset Released & Competition Live on Codabench!

The competition dataset is now released, and both tracks are live on Codabench for all phases. Submit your entries and climb the leaderboard: Track 1: Noise Event Detection โ†’  |  Track 2: Noise Event Removal โ†’

How to participate:

  1. Google Form registration has now closed (as of 9th September 2026). Teams that have already registered can continue to make submissions on Codabench.
  2. Register on Codabench for both tracks โ€” Track 1 and Track 2 โ€” to make submissions and appear on the leaderboards.
  3. Join our Discord community for discussions, updates, and support โ€” Join the Datathon Discord โ†’

Metric note (Track 2): ΔWER is now reported as a percentage (e.g. −1.17) on the leaderboard to match the official ASR scoring convention. The ranking Combined score (SI-SDR + 100×ΔWERfraction) is unchanged.

July 2026

Registration is Now Open!  Registration is Now Closed

Registration is live! Task description, dataset details, evaluation metrics, and data examples are now available. Registration is now closed (9th September 2026).

April 2026

Website Launched!

The official website for Datathon@IndoML 2026 is now live. Registration details, task description, dataset, and timeline will be announced soon.

Motivation

Most academic speech enhancement benchmarks are built around stationary noise (white, pink, cafรฉ hum) or studio-recorded mixtures. Field recordings from rural and semi-urban India contain bursty, semantically rich events โ€” a passing motorbike, a hen, a pressure-cooker whistle, a doorbell, a TV in another room โ€” that current denoisers either smear over or treat as speech.

Two consequences follow:

  • Downstream ASR fails disproportionately on speakers from these environments, amplifying the existing performance gap between high-resource and low-resource Indic languages.
  • Standard sound-event detection (SED) datasets (AudioSet, DESED, DCASE) under-represent both Indian acoustic contexts and Indian speech, making transfer learning from them brittle.

This challenge targets that gap directly. It asks participants to treat noise as a first-class, labelled, time-localised object โ€” and then to suppress it without harming the speech.

Datathon Chairs

Subhajit Datta

Dr. Subhajit Datta

Heritage Institute of Technology
Mahesh Mohan

Dr. Mahesh Mohan

IIT Kharagpur
Prasanta Kumar Ghosh

Dr. Prasanta Kumar Ghosh

IISc Bangalore
Debopriyo Banerjee

Dr. Debopriyo Banerjee

Inception - G42

Technical Volunteers

Nihar Desai

Nihar Desai

ARTPARK, IISc
Pavan Kumar J

Pavan Kumar J

ARTPARK, IISc
Sujith P

Sujith P

ARTPARK, IISc
Shubhadip Nag

Shubhadip Nag

Walmart

Competition Host (Codabench)

Shubhadip Nag

Shubhadip Nag

Walmart

Task Description

Noise Event Detection & Removal in Indic Speech

Robust, Inclusive Speech Processing for Real-World Indic Speech

Speech recordings collected in real Indian environments are dominated by non-stationary background events - vehicle horns, dogs barking, children crying, doorbells, ringtones, kitchen appliances, and devotional music. These events degrade downstream Automatic Speech Recognition (ASR).

This challenge invites participants to build a two-stage system on the Vaani dataset that (i) detects noise events with precise timestamps, and (ii) removes them while preserving the underlying speech. The challenge is framed under the Responsible AI theme, with explicit emphasis on robustness, linguistic inclusivity across multiple Indian languages, and methodological transparency.

Evaluation uses per-track metrics - F1 & Dice for detection, SI-SDR & Delta WER for removal - with PESQ evaluated for top-5 removal entries. A novelty score adjusts the final standings.

Challenge Tracks

Participants may enter either track independently or both.

01

Detection

Detect Events

Detect noise events in Indic speech recordings with precise onset/offset timestamps. Effectively utilise data annotated at different levels.

Raw Audio
Your Model
Event JSON
{onset: 1.24, offset: 3.81}, {onset: 4.31, offset: 4.71}, {onset: 5.04, offset: 5.41}
Evaluation Metrics: F1 Dice
Event timeline (onset/offset pairs) from Track 1 is passed as conditioning input to Track 2 โ€” guiding the model on where to suppress noise
02

Removal

Suppress & Preserve

Suppress the detected noise events while preserving the underlying speech signal - output clean, intelligible audio.

Audio + Events
Your Model
Clean WAV
16 kHz mono WAV - one per test clip, original filename retained
Evaluation Metrics: SI-SDR Delta WER PESQ (top-5)

Competition Process

Submission

Track 1: Submit a JSON file with onset/offset events per clip.
Track 2: Submit cleaned 16 kHz mono WAV files.

Evaluation

Automated scoring on held-out test clips. Track 1: F1 + Dice. Track 2: SI-SDR + ฮ”WER, then PESQ for top-5.

Leaderboard

Live rankings published after each submission window. Final standings adjusted by expert Novelty Score for top-5 entries.

Awards

Top teams invited to present at IndoML 2026 and receive cash prizes. Code release required for prize-eligible entries.

Evaluation Metrics

Track 1 — Detection
Primary
Event-based F1

A prediction is correct when its temporal extent overlaps with ground truth within +/-20% of event duration.

Primary
Segment Dice

Temporal overlap between predicted and reference event segments: 2 * |P intersection G| / (|P|+|G|).

Overall ranking Equal-weight average of F1 and Dice scores.
Track 2 — Removal
Primary
SI-SDR

Scale-Invariant Signal-to-Distortion Ratio between enhanced signal and synthetic clean reference.

Primary
Delta WER

A frozen multilingual Indic ASR is run on both noisy and enhanced clips. Delta WER = WERnoisy - WERenhanced. Higher is better.

Top-5 Only
PESQ

Perceptual Evaluation of Speech Quality - intelligibility & naturalness vs. clean references. Evaluated only for the top-5 initial entries.

Initial ranking Equal-weight average of SI-SDR and Delta WER. Top-5 then evaluated for PESQ.
Novelty Score - Following metric-based ranking, an expert panel applies a novelty factor to the top-5 entries in each track to determine the final standings. Rewards original contributions: new architectures, novel training regimes, principled event-conditioning, or semi-supervised approaches exploiting unlabelled Vaani audio.

Responsible AI Alignment

This challenge is positioned under the Responsible AI track on three explicit axes:

Robustness

Real-world Indic recordings - not curated studio mixtures - are the evaluation distribution. Systems are scored on actual ASR improvement (Delta WER), not signal-level metrics alone.

Inclusivity

Vaani spans multiple Indian languages and a wide range of speakers. The eval set is monitored for language-wise and class-wise balance so no sub-population is under-represented.

Transparency

Top-5 submissions must release code and document pre-trained dependencies. The frozen ASR used for Delta WER is publicly identified for independent reproducibility.

Compete

Register on the Codabench competition pages to make submissions, view the live leaderboards, and access the evaluation phase. Download the training dataset to get started.

Dataset

Vaani Corpus

A large-scale, openly released Indic speech dataset spanning multiple Indian languages, collected across districts of India. Learn more at vaani.iisc.ac.in or read the paper.

The dataset consists of three types of annotated noise events, totalling ~154.6 hours of training audio across 90,637 segments. An additional 10 hours of gold-standard audio will serve as the test set, with ground-truth timestamps held back for automated evaluation. Effectively utilising data annotated at different quality levels is a key part of the challenge.

Sample Preview on HuggingFace
# Annotation Type Duration (hrs) Segments Description
1 Verified Timestamps (๐Ÿฅ‡ Gold) 21.8 11,111 Noise events with precise timestamps where mutual agreement between multiple annotators has been verified
2 Unverified Timestamps (๐Ÿฅˆ Silver) 100.32 61,642 Annotated noise events with timestamps, but agreement between annotators is not verified
3 No Timestamps (๐Ÿฅ‰ Bronze) 32.42 17,884 Only noise event tags present in the transcript โ€” no onset/offset timestamps
Training Total 154.6 90,637
An additional 10 hours of gold-standard audio will be provided as the test set. Ground-truth timestamps are held back on the server for automated evaluation.

Important Dates

Registration (1st Aug โ€“ 9th Sep 2026)

Registration is now closed (9th September 2026)

๐Ÿš€ Competition Launch

15th August 2026

๐Ÿ“ฆ Validation Dataset Release

4th September 2026 (released)

๐ŸŽ™๏ธ AMA Session #1

9th September 2026 (held) โ€” recording ยท slides

๐Ÿ Phase 1 (Half-Marathon) Results

20th September 2026 Extended to 22nd September 2026 (due to Codabench downtime) โ€” Top teams selected for half-marathon prizes (both tracks). Track 2 material submission (Google Form) closes 24th September 2026, 11:59 PM IST.

Competition End / Final Submission Deadline

15th October 2026 Extended to 17th October 2026 (2-day extension, due to Codabench downtime)

Results Announcement

TBA

Presentation at IndoML 2026

18โ€“21 December 2026

All deadlines will be at 12:00 Noon IST (Indian Standard Time).

Prizes

Cash Prizes

A total prize pool of โ‚น2,00,000 โ€” โ‚น1,00,000 (1 lakh) per track, awarded across the top teams in each track.

Present at IndoML 2026

Top teams will be invited to present their solutions at IndoML 2026, in front of leading researchers from academia and industry.

Novelty Award

An expert panel award for the most original methodological contribution: new architectures, novel training regimes, or unsupervised approaches.

Prize Money Distribution

Total budget: โ‚น2,00,000. Distributed across Half-Marathon, Final Stage, and Spotlight awards.

Phase 1 โ€” Half-Marathon Prizes (โ‚น40,000)

Awarded on 22nd September 2026 (extended from 20th September due to Codabench downtime) based on leaderboard standings at the midpoint. Note: The leaderboard will not be reset after the half-marathon.

CategoryPer TeamTrack 1Track 2Total
๐Ÿฅ‡๐Ÿฅˆ๐Ÿฅ‰ Ranks 1โ€“3โ‚น5,000โ‚น15,000โ‚น15,000โ‚น30,000
Ranks 4โ€“8โ‚น1,000โ‚น5,000โ‚น5,000โ‚น10,000
Half-Marathon Totalโ‚น40,000

Final Stage Prizes (โ‚น1,50,000)

Ranking: Leaderboard (80%) + Novelty (15%) + Report (5%)

CategoryTrack 1Track 2Total
๐Ÿฅ‡ 1st Placeโ‚น35,000โ‚น35,000โ‚น70,000
๐Ÿฅˆ 2nd Placeโ‚น15,000โ‚น15,000โ‚น30,000
๐Ÿฅ‰ 3rd Placeโ‚น10,000โ‚น10,000โ‚น20,000
4thโ€“8th Place (โ‚น3,000 each)โ‚น15,000โ‚น15,000โ‚น30,000
Final Stage Totalโ‚น1,50,000
Spotlight Solution Award โ€” โ‚น10,000 (Track 2 only): An additional โ‚น10,000 reserved exclusively for Track 2 (Noise Event Removal) as a spotlight prize for exceptional solution(s), given that it is the harder problem.
Grand Total Prize Pool: โ‚น2,00,000

Half-Marathon โ‚น40K + Final Stage โ‚น1.5L + Spotlight โ‚น10K (Track 2)

Top-5 entries per track must release training and inference code under a permissive open-source licence and submit a short (<=4 pages) system description.

FAQ

Who can participate?

Anyone based in India can participate โ€” students, researchers, and industry professionals are all welcome. Each team must include at least one member affiliated with an Indian university, research institution, or organisation based in India.

Is there a team size limit?

There is no strict cap, but based on previous editions the sweet spot is 3โ€“4 members. Slightly larger teams are fine, though very large teams are discouraged. Each participant may only join one team.

Can I participate in just one of the two tracks?

Yes, there is no restriction. You can participate in one or both tracks โ€” it's entirely up to you.

What is Phase 1 (Half-Marathon)? Do only selected teams proceed to Phase 2?

No, everyone proceeds to Phase 2. Phase 1 (Half-Marathon) is an intermediate checkpoint with its own prizes, like a half-marathon before the full marathon. The leaderboard is not reset โ€” all teams continue into Phase 2 regardless of their Phase 1 standing.

How is the leaderboard ranked?

The leaderboard is ranked by the Combined score. For Track 1, this is F1 + Dice. For Track 2, this is SI-SDR + 100 ร— ฮ”WER (fraction). Individual metric scores are shown for transparency, but the Combined score determines your rank.

Where can I find the dataset and problem statement?

Go to the Codabench competition page (Track 1 or Track 2), then open "Get Started" โ†’ "Files". Download input_data there (the validation dataset is available in this section for both tracks). Training data is available on Hugging Face.

What is the last date to register?

Registration is now closed โ€” Google Form registration closed on 9th September 2026 at 12:00 PM IST. Teams that have already registered can continue to make submissions on Codabench.

What models and data can we use?

We follow an open model and open data policy. Teams may use any publicly available, closed-source, or proprietary models, along with additional data or augmentation strategies.

How will submissions be evaluated?

Track 1 (Detection): Combined = Event-based F1 + Segment Dice. Track 2 (Removal): Combined = SI-SDR + 100 ร— ฮ”WER. Top-5 entries in Track 2 are additionally evaluated on PESQ. Final standings for top-5 in each track are adjusted by an expert Novelty Score.

Will top teams get travel support?

Top-performing teams will be invited to present at IndoML 2026. Details regarding travel support will be communicated later.

Contact Us

Expected Outcomes

Public Benchmark

A public benchmark for noise-event-aware speech enhancement on Indic audio โ€” a gap that currently has no widely adopted dataset.

Open-Source Systems

Open-sourced winning systems, raising the floor of available denoising tools for Indian-language ASR.

Evaluation Harness

A reusable evaluation harness pairing event-detection metrics with downstream ฮ”WER on real audio.

Previous Editions