Results status · public leaderboard

Leaderboard.

Warm-up phase, live. 194 teams across four tracks, best submission per team. Refreshed daily from Codabench, last Oct 10, 2026 12:11 UTC. Final positions are gated by reproducibility audit.

EEG-to-Image · Track 01.

Rank held-out candidate images from a single EEG epoch. Targets are frozen DINOv2-giant embeddings, so the score isolates the EEG side. Controlled shift: test stimuli are unseen during training. Higher is better.

MetricTop-5 retrieval accuracy
Tie-breakTop-1 accuracy
SponsorAlljoined
Track 1 warm-up leaderboard. Top 10 of 50 teams, best submission per team. Full board on Codabench ↗
RankTeamModelTop-5 accuracyGap to leaderSubmitted
1louisliu—0.58—Oct 09
2evo-brain—0.550.03Oct 10
3Shenzhen Campus of Sun Yat-sen University—0.520.06Oct 05
4ponpopon—0.520.06Oct 04
5abe_lincoln—0.510.07Oct 04
6_ares_—0.500.08Oct 04
7Ethan—0.480.10Oct 10
8cchen847—0.470.11Oct 03
9rohit_kumar_varma—0.470.11Oct 08
10cheesypasta—0.460.12Oct 07
50 teams · leader louisliu at 0.58 · updated Oct 10, 2026 12:11 UTC How final scoring works ↓

BCI decoding · Track 02.

Predict the cued command (motor imagery, mental math, word association) on later sessions of the same subject, with no per-session recalibration. The score reflects calibration-free stability across session drift. Higher is better.

MetricBalanced accuracy
Tie-breakMean per-class F1
SponsorMeta FAIR Brain & AI
Track 2 warm-up leaderboard. Top 10 of 50 teams, best submission per team. Full board on Codabench ↗
RankTeamModelBalanced accuracyGap to leaderSubmitted
1cheesypasta—0.94—Oct 10
2Ethan—0.940.00Oct 09
3speedwagon—0.930.01Oct 06
4天色天歌天籁音CleanVM-44 (zero test-pool labels)0.930.01Oct 06
5shamadei—0.920.02Sep 28
6jiahuian1—0.920.02Oct 05
7sjhan—0.920.02Oct 04
8abe_lincoln—0.920.02Oct 02
9period—0.910.03Oct 04
10h5371hCao-NI-MAE (https://doi.org/10.1145/3774906.3802770)0.910.03Sep 25
50 teams · leader cheesypasta at 0.94 · updated Oct 10, 2026 12:11 UTC How scoring works ↓

Sleep onset · Track 03.

From each point in a continuous wearable EEG recording, predict the seconds remaining until the first N2 epoch. Models are evaluated on new recordings from participants represented in training and on recordings from completely unseen participants. Weighted errors emphasize predictions closer to sleep onset, and the final score macro-averages performance across the seen- and unseen-subject groups. Lower is better.

MetricMacro W-bMAE (s)
GroupsSeen + unseen subjects
SponsorMuse
Track 3 warm-up leaderboard. Top 10 of 50 teams, best submission per team. Full board on Codabench ↗
RankTeamModelWeighted bMAE (s)Gap to leaderSubmitted
1behradbeyglo—27.88—Oct 08
2Ethan—45.4817.60Oct 10
3PurvKabaria—54.2126.33Oct 09
4abe_lincoln—59.1731.29Oct 08
5ahmadabdelaal—66.9539.07Oct 02
6wooluck98—75.4247.54Oct 01
7MannasAI—83.3055.42Sep 29
8LIN-BTU-AIDuMa-Sleep83.4455.56Sep 23
9mehular—86.4258.54Oct 09
10fmashrur—95.7367.85Sep 20
50 teams · leader behradbeyglo at 27.88 · updated Oct 10, 2026 12:11 UTC How scoring works ↓

EMG-to-Pose · Track 04.

Regress 20-joint angle trajectories from 16-channel wrist surface EMG. The cross-user split varies anatomy and device placement, so the score rewards robust pose representations rather than participant-specific templates. Lower is better.

MetricMean absolute angular error (degrees)
Tie-breakFinal protocol pending
SponsorMeta Reality Labs
Track 4 warm-up leaderboard. Top 10 of 44 teams, best submission per team. Full board on Codabench ↗
RankTeamModelAngular MAE (°)Gap to leaderSubmitted
1Ethan—10.59—Oct 10
2tom_ml—10.740.15Oct 04
3sjhan—10.980.39Oct 07
4evo-brain—11.190.60Oct 09
5makoshan—11.230.64Oct 07
6CortixAI—11.540.95Oct 09
7Rajat—12.191.60Oct 09
8Team Yuyu—12.211.62Oct 06
9Young—12.612.02Oct 10
10BNEL—12.762.17Oct 01
44 teams · leader Ethan at 10.59 · updated Oct 10, 2026 12:11 UTC How scoring works ↓

How a sealed-phase submission becomes a final leaderboard score.

These rules describe final ranking on the private 2026 evaluation data. Warm-up uses public data and may use a proxy task or metric, as documented on each Codabench track page.

Open scoring methodology
  1. Official results CODABENCH

    Codabench is authoritative.

    Codabench queues each upload, evaluates it when a worker is available, and publishes the official score. This page summarises the ranking rules. Use your track's Codabench leaderboard for live results.

    Final phaseOct 28 to Nov 21, 2026
    Daily cap5/day warm-up · 1/day sealed final
  2. Aggregation BEST-OF-5

    Final score = best of last five.

    The public board shows your best-ever number. The final workshop ranking, however, only considers your last five submissions. This rewards focused iteration over exhaustive lottery search.

    Public boardBest ever
    Final rankingBest of last 5
  3. Reproducibility audit NOV 21

    Top-3 per track replay from config.

    We re-run the committed training pipeline against the sealed split. Within ±2 σ of the submitted score, you stay on the board. Outside, you drop. Audit is led by Arnaud Delorme (EEGLAB).

    Tolerance±2 σ on metric
    Audit windowAfter Nov 21

Planned statistical analysis, mirrored from the proposal.

Official track rankings use the point metric published on Codabench. The equations below describe the complementary uncertainty, rank-stability, and cross-track analyses planned after evaluation. They do not replace the track ranking metric.

Open formal definitions and code

Prediction, unit score, and track score

A submission \(a\) receives a hidden signal \(X_{t,i}\) and metadata \(m_{t,i}\) for track \(t\), then writes a prediction \(\hat{y}_{a,t,i}\). Codabench keeps \(y_{t,i}\) hidden and computes the official point score.

\[ \hat{y}_{a,t,i} = f_a(X_{t,i}, m_{t,i}) \]

Examples are first collapsed into independent bootstrap units \(u \in U_t\): subject-image query blocks for EEG-to-IMG, subject-session-context cells for BCI, recordings for sleep onset, and participant-stage blocks for EMG-to-Pose. Each unit gets an oriented contribution \(s_{a,t,u}\), where higher is always better. For MAE we use the negative error internally.

\[ s_{a,t,u} = \mathrm{score}_t(\hat{y}_{a,t,u}, y_{t,u}) \] \[ \mathcal{S}_{a,t} = \frac{1}{|U_t|}\sum_{u \in U_t} s_{a,t,u} \]

The visible leaderboard for track \(t\) is the point-estimate ordering of \(\mathcal{S}_{a,t}\).

python · predictions to S_a,t
1import numpy as np
2
3def build_unit_scores(y_pred, y_true, unit_ids, score_unit, lower_is_better=False):
4 """Return s_{a,t,u} after collapsing examples into units."""
5 y_pred, y_true, unit_ids = map(np.asarray, (y_pred, y_true, unit_ids))
6 scores = []
7 for unit in np.unique(unit_ids):
8 idx = unit_ids == unit
9 value = score_unit(y_pred[idx], y_true[idx])
10 scores.append(-value if lower_is_better else value)
11 return np.asarray(scores, dtype=float)
12
13def track_score(unit_scores):
14 """Compute S_{a,t}. All returned scores are higher-is-better."""
15 return float(np.mean(unit_scores))

Confidence interval, p-value, and rank stability

For bootstrap draw \(b\), the post-evaluation analysis resamples independent units \(U_t^{(b)}\) and recomputes each team's score. Pairwise uncertainty is calculated on the paired score difference, not from two separate confidence intervals.

\[ \mathcal{S}^{(b)}_{a,t} = \frac{1}{|U_t^{(b)}|}\sum_{u \in U_t^{(b)}} s_{a,t,u} \] \[ \Delta^{(b)}_{a,c,t} = \mathcal{S}^{(b)}_{a,t} - \mathcal{S}^{(b)}_{c,t} \] \[ \begin{aligned} \mathrm{CI}_{95}(\Delta_{a,c,t}) = \big[&q_{0.025}(\Delta^{(b)}_{a,c,t}),\\ &q_{0.975}(\Delta^{(b)}_{a,c,t})\big] \end{aligned} \]

If this interval contains zero, neighbouring teams are flagged as statistically indistinguishable. For prize-relevant comparisons, the two-sided bootstrap p-value is Holm-adjusted.

\[ \begin{aligned} p_{\mathrm{boot}} = 2\min\big(&\Pr_b[\Delta^{(b)}_{a,c,t} \le 0],\\ &\Pr_b[\Delta^{(b)}_{a,c,t} \ge 0]\big) \end{aligned} \]

Rank stability is computed by re-ranking all teams inside each bootstrap draw.

\[ r^{(b)}_{a,t} = \mathrm{rank}\left(\mathcal{S}^{(b)}_{a,t}\right) \] \[ \begin{gathered} \Pr(r^{(b)}_{a,t}\le1),\\ \Pr(r^{(b)}_{a,t}\le3),\\ \Pr(r^{(b)}_{a,t}\le5) \end{gathered} \]
python · CI, p_boot, Holm, ranks
1import numpy as np
2from confidence_intervals import get_bootstrap_indices, get_conf_int
3from statsmodels.stats.multitest import multipletests
4
5def bootstrap_track(unit_scores, unit_ids=None, n_boot=10_000):
6 teams = list(unit_scores)
7 n_units = len(unit_scores[teams[0]])
8 score_boot = {team: np.empty(n_boot) for team in teams}
9 rank_boot = {team: np.empty(n_boot, dtype=int) for team in teams}
10 for b in range(n_boot):
11 idx = get_bootstrap_indices(n_units, conditions=unit_ids, random_state=b)
12 scores = {team: float(np.mean(np.asarray(vals)[idx])) for team, vals in unit_scores.items()}
13 for rank, team in enumerate(sorted(teams, key=scores.get, reverse=True), start=1):
14 score_boot[team][b] = scores[team]
15 rank_boot[team][b] = rank
16 return score_boot, rank_boot
17
18def pair_summary(score_boot, team_a, team_c, alpha=5):
19 delta = score_boot[team_a] - score_boot[team_c]
20 ci_low, ci_high = get_conf_int(delta, alpha=alpha)
21 p_boot = 2 * min(np.mean(delta <= 0), np.mean(delta >= 0))
22 return {"delta": float(np.mean(delta)), "ci95": (float(ci_low), float(ci_high)),
23 "p_boot": min(float(p_boot), 1.0), "indistinguishable": bool(ci_low <= 0 <= ci_high)}
24
25def add_holm(rows, alpha=0.05):
26 reject, p_holm, _, _ = multipletests([r["p_boot"] for r in rows], method="holm", alpha=alpha)
27 for row, adj_p, keep in zip(rows, p_holm, reject):
28 row.update(p_holm=float(adj_p), significant_after_holm=bool(keep))
29 return rows
30
31def rank_stability(rank_boot):
32 return {team: {"top1": float(np.mean(r <= 1)), "top3": float(np.mean(r <= 3)),
33 "top5": float(np.mean(r <= 5))} for team, r in rank_boot.items()}

Test-set sizing

The hidden-test size for each track is chosen so the expected half-width of the 95% interval falls below \(\nu_t\), the smallest practically meaningful difference for that track. \(\hat{\sigma}_t\) is the pilot standard deviation at the top-level bootstrap unit and \(n_{\mathrm{eff},t}\) is the number of independent held-out units. If a dataset cannot support this target, intervals widen and ties are reported rather than over-interpreting small margins.

\[ 1.96\,\hat{\sigma}_t / \sqrt{n_{\mathrm{eff},t}} \le \nu_t \]
python · test-set sizing
1import math
2
3def ci_half_width(sigma_hat, n_eff, z=1.96):
4 return z * sigma_hat / math.sqrt(n_eff)
5
6def required_n_eff(sigma_hat, nu_t, z=1.96):
7 """Smallest independent hidden-test count satisfying the target half-width."""
8 return math.ceil((z * sigma_hat / nu_t) ** 2)
9
10def meets_resolution_target(sigma_hat, n_eff, nu_t):
11 return ci_half_width(sigma_hat, n_eff) <= nu_t

Overall ranking

Each valid submission gets rank points \(P_{\mathrm{team},t}\) on its track (linearly interpolated against the field, so the top of the field scores 1 and the bottom scores 0). The submitted-track average summarises a team's record across the tracks it entered. The all-track score averages over all four task-specific tracks, padding missing tracks with zero so transfer is rewarded over single-track wins. \(r_{\mathrm{team},t}\) is the team's rank, \(N_t\) is the number of valid submissions on the track, and \(T_{\mathrm{team}}\) is the set of tracks the team submitted.

\[ P_{\mathrm{team},t} = \begin{cases} 1-\dfrac{r_{\mathrm{team},t}-1}{N_t-1}, & N_t>1, \\ 1, & N_t=1 \end{cases} \] \[ \mathcal{S}_{\mathrm{submitted}}(\mathrm{team}) = \frac{1}{|T_{\mathrm{team}}|}\sum_{t\in T_{\mathrm{team}}} P_{\mathrm{team},t} \] \[ \mathcal{S}_{\mathrm{all}}(\mathrm{team}) = \frac{1}{4}\sum_{t\in T} P^{\star}_{\mathrm{team},t} \] \[ P^{\star}_{\mathrm{team},t} = \begin{cases} P_{\mathrm{team},t}, & t\in T_{\mathrm{team}}, \\ 0, & t\notin T_{\mathrm{team}} \end{cases} \]
python · rank-point aggregation
1def rank_points(rank, n_submissions):
2 return 1.0 if n_submissions == 1 else 1.0 - (rank - 1) / (n_submissions - 1)
3
4def submitted_track_score(points_by_track, submitted_tracks):
5 return sum(points_by_track[t] for t in submitted_tracks) / len(submitted_tracks)
6
7def all_track_score(points_by_track, all_tracks=("img", "bci", "sleep", "emg")):
8 """Missing tracks get zero points."""
9 return sum(points_by_track.get(t, 0.0) for t in all_tracks) / len(all_tracks)

Take a baseline and beat it.

The competition runs Sep 21 to Nov 21, 2026. Prepare in NeuralBench, reproduce a track baseline, and use the warm-up phase to iterate before final evaluation.