Leaderboard.
Warm-up phase, live. 194 teams across four tracks, best submission per team. Refreshed daily from Codabench, last Oct 10, 2026 12:11 UTC. Final positions are gated by reproducibility audit.
EEG-to-Image · Track 01.
Rank held-out candidate images from a single EEG epoch. Targets are frozen DINOv2-giant embeddings, so the score isolates the EEG side. Controlled shift: test stimuli are unseen during training. Higher is better.
| Rank | Team | Model | Top-5 accuracy | Gap to leader | Submitted |
|---|---|---|---|---|---|
| 1 | louisliu | — | 0.58 | — | Oct 09 |
| 2 | evo-brain | — | 0.55 | 0.03 | Oct 10 |
| 3 | Shenzhen Campus of Sun Yat-sen University | — | 0.52 | 0.06 | Oct 05 |
| 4 | ponpopon | — | 0.52 | 0.06 | Oct 04 |
| 5 | abe_lincoln | — | 0.51 | 0.07 | Oct 04 |
| 6 | _ares_ | — | 0.50 | 0.08 | Oct 04 |
| 7 | Ethan | — | 0.48 | 0.10 | Oct 10 |
| 8 | cchen847 | — | 0.47 | 0.11 | Oct 03 |
| 9 | rohit_kumar_varma | — | 0.47 | 0.11 | Oct 08 |
| 10 | cheesypasta | — | 0.46 | 0.12 | Oct 07 |
BCI decoding · Track 02.
Predict the cued command (motor imagery, mental math, word association) on later sessions of the same subject, with no per-session recalibration. The score reflects calibration-free stability across session drift. Higher is better.
| Rank | Team | Model | Balanced accuracy | Gap to leader | Submitted |
|---|---|---|---|---|---|
| 1 | cheesypasta | — | 0.94 | — | Oct 10 |
| 2 | Ethan | — | 0.94 | 0.00 | Oct 09 |
| 3 | speedwagon | — | 0.93 | 0.01 | Oct 06 |
| 4 | 天色天歌天籁音 | CleanVM-44 (zero test-pool labels) | 0.93 | 0.01 | Oct 06 |
| 5 | shamadei | — | 0.92 | 0.02 | Sep 28 |
| 6 | jiahuian1 | — | 0.92 | 0.02 | Oct 05 |
| 7 | sjhan | — | 0.92 | 0.02 | Oct 04 |
| 8 | abe_lincoln | — | 0.92 | 0.02 | Oct 02 |
| 9 | period | — | 0.91 | 0.03 | Oct 04 |
| 10 | h5371h | Cao-NI-MAE (https://doi.org/10.1145/3774906.3802770) | 0.91 | 0.03 | Sep 25 |
Sleep onset · Track 03.
From each point in a continuous wearable EEG recording, predict the seconds remaining until the first N2 epoch. Models are evaluated on new recordings from participants represented in training and on recordings from completely unseen participants. Weighted errors emphasize predictions closer to sleep onset, and the final score macro-averages performance across the seen- and unseen-subject groups. Lower is better.
| Rank | Team | Model | Weighted bMAE (s) | Gap to leader | Submitted |
|---|---|---|---|---|---|
| 1 | behradbeyglo | — | 27.88 | — | Oct 08 |
| 2 | Ethan | — | 45.48 | 17.60 | Oct 10 |
| 3 | PurvKabaria | — | 54.21 | 26.33 | Oct 09 |
| 4 | abe_lincoln | — | 59.17 | 31.29 | Oct 08 |
| 5 | ahmadabdelaal | — | 66.95 | 39.07 | Oct 02 |
| 6 | wooluck98 | — | 75.42 | 47.54 | Oct 01 |
| 7 | MannasAI | — | 83.30 | 55.42 | Sep 29 |
| 8 | LIN-BTU-AI | DuMa-Sleep | 83.44 | 55.56 | Sep 23 |
| 9 | mehular | — | 86.42 | 58.54 | Oct 09 |
| 10 | fmashrur | — | 95.73 | 67.85 | Sep 20 |
EMG-to-Pose · Track 04.
Regress 20-joint angle trajectories from 16-channel wrist surface EMG. The cross-user split varies anatomy and device placement, so the score rewards robust pose representations rather than participant-specific templates. Lower is better.
| Rank | Team | Model | Angular MAE (°) | Gap to leader | Submitted |
|---|---|---|---|---|---|
| 1 | Ethan | — | 10.59 | — | Oct 10 |
| 2 | tom_ml | — | 10.74 | 0.15 | Oct 04 |
| 3 | sjhan | — | 10.98 | 0.39 | Oct 07 |
| 4 | evo-brain | — | 11.19 | 0.60 | Oct 09 |
| 5 | makoshan | — | 11.23 | 0.64 | Oct 07 |
| 6 | CortixAI | — | 11.54 | 0.95 | Oct 09 |
| 7 | Rajat | — | 12.19 | 1.60 | Oct 09 |
| 8 | Team Yuyu | — | 12.21 | 1.62 | Oct 06 |
| 9 | Young | — | 12.61 | 2.02 | Oct 10 |
| 10 | BNEL | — | 12.76 | 2.17 | Oct 01 |
How a sealed-phase submission becomes a final leaderboard score.
These rules describe final ranking on the private 2026 evaluation data. Warm-up uses public data and may use a proxy task or metric, as documented on each Codabench track page.
Open scoring methodology
-
Official results CODABENCH
Codabench is authoritative.
Codabench queues each upload, evaluates it when a worker is available, and publishes the official score. This page summarises the ranking rules. Use your track's Codabench leaderboard for live results.
Final phaseOct 28 to Nov 21, 2026Daily cap5/day warm-up · 1/day sealed final -
Aggregation BEST-OF-5
Final score = best of last five.
The public board shows your best-ever number. The final workshop ranking, however, only considers your last five submissions. This rewards focused iteration over exhaustive lottery search.
Public boardBest everFinal rankingBest of last 5 -
Reproducibility audit NOV 21
Top-3 per track replay from config.
We re-run the committed training pipeline against the sealed split. Within ±2 σ of the submitted score, you stay on the board. Outside, you drop. Audit is led by Arnaud Delorme (EEGLAB).
Tolerance±2 σ on metricAudit windowAfter Nov 21
Planned statistical analysis, mirrored from the proposal.
Official track rankings use the point metric published on Codabench. The equations below describe the complementary uncertainty, rank-stability, and cross-track analyses planned after evaluation. They do not replace the track ranking metric.
Open formal definitions and code
Prediction, unit score, and track score
A submission \(a\) receives a hidden signal \(X_{t,i}\) and metadata \(m_{t,i}\) for track \(t\), then writes a prediction \(\hat{y}_{a,t,i}\). Codabench keeps \(y_{t,i}\) hidden and computes the official point score.
\[ \hat{y}_{a,t,i} = f_a(X_{t,i}, m_{t,i}) \]Examples are first collapsed into independent bootstrap units \(u \in U_t\): subject-image query blocks for EEG-to-IMG, subject-session-context cells for BCI, recordings for sleep onset, and participant-stage blocks for EMG-to-Pose. Each unit gets an oriented contribution \(s_{a,t,u}\), where higher is always better. For MAE we use the negative error internally.
\[ s_{a,t,u} = \mathrm{score}_t(\hat{y}_{a,t,u}, y_{t,u}) \] \[ \mathcal{S}_{a,t} = \frac{1}{|U_t|}\sum_{u \in U_t} s_{a,t,u} \]The visible leaderboard for track \(t\) is the point-estimate ordering of \(\mathcal{S}_{a,t}\).
1import numpy as np23def build_unit_scores(y_pred, y_true, unit_ids, score_unit, lower_is_better=False):4"""Return s_{a,t,u} after collapsing examples into units."""5y_pred, y_true, unit_ids = map(np.asarray, (y_pred, y_true, unit_ids))6scores = []7for unit in np.unique(unit_ids):8idx = unit_ids == unit9value = score_unit(y_pred[idx], y_true[idx])10scores.append(-value if lower_is_better else value)11return np.asarray(scores, dtype=float)1213def track_score(unit_scores):14"""Compute S_{a,t}. All returned scores are higher-is-better."""15return float(np.mean(unit_scores))
Confidence interval, p-value, and rank stability
For bootstrap draw \(b\), the post-evaluation analysis resamples independent units \(U_t^{(b)}\) and recomputes each team's score. Pairwise uncertainty is calculated on the paired score difference, not from two separate confidence intervals.
\[ \mathcal{S}^{(b)}_{a,t} = \frac{1}{|U_t^{(b)}|}\sum_{u \in U_t^{(b)}} s_{a,t,u} \] \[ \Delta^{(b)}_{a,c,t} = \mathcal{S}^{(b)}_{a,t} - \mathcal{S}^{(b)}_{c,t} \] \[ \begin{aligned} \mathrm{CI}_{95}(\Delta_{a,c,t}) = \big[&q_{0.025}(\Delta^{(b)}_{a,c,t}),\\ &q_{0.975}(\Delta^{(b)}_{a,c,t})\big] \end{aligned} \]If this interval contains zero, neighbouring teams are flagged as statistically indistinguishable. For prize-relevant comparisons, the two-sided bootstrap p-value is Holm-adjusted.
\[ \begin{aligned} p_{\mathrm{boot}} = 2\min\big(&\Pr_b[\Delta^{(b)}_{a,c,t} \le 0],\\ &\Pr_b[\Delta^{(b)}_{a,c,t} \ge 0]\big) \end{aligned} \]Rank stability is computed by re-ranking all teams inside each bootstrap draw.
\[ r^{(b)}_{a,t} = \mathrm{rank}\left(\mathcal{S}^{(b)}_{a,t}\right) \] \[ \begin{gathered} \Pr(r^{(b)}_{a,t}\le1),\\ \Pr(r^{(b)}_{a,t}\le3),\\ \Pr(r^{(b)}_{a,t}\le5) \end{gathered} \]1import numpy as np2from confidence_intervals import get_bootstrap_indices, get_conf_int3from statsmodels.stats.multitest import multipletests45def bootstrap_track(unit_scores, unit_ids=None, n_boot=10_000):6teams = list(unit_scores)7n_units = len(unit_scores[teams[0]])8score_boot = {team: np.empty(n_boot) for team in teams}9rank_boot = {team: np.empty(n_boot, dtype=int) for team in teams}10for b in range(n_boot):11idx = get_bootstrap_indices(n_units, conditions=unit_ids, random_state=b)12scores = {team: float(np.mean(np.asarray(vals)[idx])) for team, vals in unit_scores.items()}13for rank, team in enumerate(sorted(teams, key=scores.get, reverse=True), start=1):14score_boot[team][b] = scores[team]15rank_boot[team][b] = rank16return score_boot, rank_boot1718def pair_summary(score_boot, team_a, team_c, alpha=5):19delta = score_boot[team_a] - score_boot[team_c]20ci_low, ci_high = get_conf_int(delta, alpha=alpha)21p_boot = 2 * min(np.mean(delta <= 0), np.mean(delta >= 0))22return {"delta": float(np.mean(delta)), "ci95": (float(ci_low), float(ci_high)),23"p_boot": min(float(p_boot), 1.0), "indistinguishable": bool(ci_low <= 0 <= ci_high)}2425def add_holm(rows, alpha=0.05):26reject, p_holm, _, _ = multipletests([r["p_boot"] for r in rows], method="holm", alpha=alpha)27for row, adj_p, keep in zip(rows, p_holm, reject):28row.update(p_holm=float(adj_p), significant_after_holm=bool(keep))29return rows3031def rank_stability(rank_boot):32return {team: {"top1": float(np.mean(r <= 1)), "top3": float(np.mean(r <= 3)),33"top5": float(np.mean(r <= 5))} for team, r in rank_boot.items()}
Test-set sizing
The hidden-test size for each track is chosen so the expected half-width of the 95% interval falls below \(\nu_t\), the smallest practically meaningful difference for that track. \(\hat{\sigma}_t\) is the pilot standard deviation at the top-level bootstrap unit and \(n_{\mathrm{eff},t}\) is the number of independent held-out units. If a dataset cannot support this target, intervals widen and ties are reported rather than over-interpreting small margins.
\[ 1.96\,\hat{\sigma}_t / \sqrt{n_{\mathrm{eff},t}} \le \nu_t \]1import math23def ci_half_width(sigma_hat, n_eff, z=1.96):4return z * sigma_hat / math.sqrt(n_eff)56def required_n_eff(sigma_hat, nu_t, z=1.96):7"""Smallest independent hidden-test count satisfying the target half-width."""8return math.ceil((z * sigma_hat / nu_t) ** 2)910def meets_resolution_target(sigma_hat, n_eff, nu_t):11return ci_half_width(sigma_hat, n_eff) <= nu_t
Overall ranking
Each valid submission gets rank points \(P_{\mathrm{team},t}\) on its track (linearly interpolated against the field, so the top of the field scores 1 and the bottom scores 0). The submitted-track average summarises a team's record across the tracks it entered. The all-track score averages over all four task-specific tracks, padding missing tracks with zero so transfer is rewarded over single-track wins. \(r_{\mathrm{team},t}\) is the team's rank, \(N_t\) is the number of valid submissions on the track, and \(T_{\mathrm{team}}\) is the set of tracks the team submitted.
\[ P_{\mathrm{team},t} = \begin{cases} 1-\dfrac{r_{\mathrm{team},t}-1}{N_t-1}, & N_t>1, \\ 1, & N_t=1 \end{cases} \] \[ \mathcal{S}_{\mathrm{submitted}}(\mathrm{team}) = \frac{1}{|T_{\mathrm{team}}|}\sum_{t\in T_{\mathrm{team}}} P_{\mathrm{team},t} \] \[ \mathcal{S}_{\mathrm{all}}(\mathrm{team}) = \frac{1}{4}\sum_{t\in T} P^{\star}_{\mathrm{team},t} \] \[ P^{\star}_{\mathrm{team},t} = \begin{cases} P_{\mathrm{team},t}, & t\in T_{\mathrm{team}}, \\ 0, & t\notin T_{\mathrm{team}} \end{cases} \]1def rank_points(rank, n_submissions):2return 1.0 if n_submissions == 1 else 1.0 - (rank - 1) / (n_submissions - 1)34def submitted_track_score(points_by_track, submitted_tracks):5return sum(points_by_track[t] for t in submitted_tracks) / len(submitted_tracks)67def all_track_score(points_by_track, all_tracks=("img", "bci", "sleep", "emg")):8"""Missing tracks get zero points."""9return sum(points_by_track.get(t, 0.0) for t in all_tracks) / len(all_tracks)
Take a baseline and beat it.
The competition runs Sep 21 to Nov 21, 2026. Prepare in NeuralBench, reproduce a track baseline, and use the warm-up phase to iterate before final evaluation.