Final Results

Results & Leaderboard

Final official standings for LLM as a Judge?: From Statute Prediction to Sycophancy Detection in Law — the authoritative source for track rankings.

Final · Verified

Final verification of the results is complete. The standings below are the final evaluation results. No further result submissions or corrections will be accepted. This page is the authoritative source for ranking and leaderboard information.


§ 01 · OVERVIEW

Overview

Following the initial release of the Task 1 and Task 2 results, the organizing team reviewed every request for clarification regarding the leaderboard, evaluation results, and submission formatting, and applied corrections wherever applicable. Throughout, the objective was to ensure that no participant was excluded solely because of technical or formatting issues in their submission.

17
Teams · Task 1
8
Teams · Task 2
6
Task 1 Metrics
15
Sept · Notes Due

§ 02 · TASK 1 LEADERBOARD

Task 1 — Explainable Statute Prediction

The final Task 1 leaderboard reports the best-performing run for each team, together with the corresponding evaluation metrics and Total score.

1st
KLH
0.48215
Total · Run 2
2nd
DwaipayanDatta
0.4330
Total · Run 2
3rd
AnastasiiaPotiagalova
0.3865
Total · Run 1
# Team Best Run Macro-F1 Micro-F1 Accuracy ROUGE-L BLEU METEOR Total
1 KLH Run 2 0.69090.79130.54380.34170.20560.3196 0.48215
2 DwaipayanDatta Run 2 0.66750.80000.66660.19610.06010.2075 0.4330
3 AnastasiiaPotiagalova Run 1 0.63460.72150.36840.23850.10320.2527 0.3865
4 PraptipriyaPhukon Run 2 0.55000.66180.52630.18190.06850.1931 0.3636
5 SayanjibSur Run 1 0.45920.64060.47360.16600.08170.1820 0.3339
6 TokenX Run 1 0.49410.58590.38590.19500.09230.2422 0.3326
7 Sruthy Run 1 0.48530.63560.45610.14650.05610.1419 0.3203
8 TanishqKathar Run 1 0.53540.61840.40350.15020.05070.1603 0.3198
9 KunjanKalita Run 1 0.48370.59110.36840.18650.08390.1789 0.3154
10 AnjaliR Run 1 0.44520.53250.33330.20610.10590.1954 0.3031
11 Dasa Sai Shankar Run 2 0.55060.62420.36840.05420.02540.1294 0.2920
12 PrajaktaBodakhe Run 1 0.52980.55030.33330.12780.03330.1347 0.2849
13 PrasenjitRoy Run 1 0.43310.59610.42100.10150.03470.0874 0.2790
14 Pavithira Run 2 0.41890.52330.36840.12330.05440.1282 0.2694
15 NithinKumarHeraje Run 1 0.38880.61680.42100.07770.02040.0594 0.2640
16 HarshSingh Run 1/2 0.44470.53080.28070.09770.02250.0794 0.2426
17 Mirnmoy Run 1 0.26920.38630.05260.18360.10030.1700 0.1937
Scoring

For Task 1, equal weightage was assigned to Macro-F1, Micro-F1, Accuracy, ROUGE-L, BLEU, and METEOR. The Total score is the arithmetic mean of these six metrics.


§ 03 · TASK 2 LEADERBOARD

Task 2 — Sycophancy Detection

The Task 2 results are unchanged from the previously communicated results. There were no new submissions or formatting-related corrections requiring a change to the Task 2 evaluation. Teams are ranked by Macro-F1.

1st
AnastasiiaPotiagalova
81.29
Macro-F1 % · Run 1
2nd
RadhikaBohra
76.81
Macro-F1 % · Run 2
3rd
SupriyaChanda
65.04
Macro-F1 % · Run 2
# Team Best Run Accuracy % Precision % Recall % F1 (Syc) % Macro-F1 %
1 AnastasiiaPotiagalova Run 1 84.3365.3584.6273.74 81.29
2 RadhikaBohra Run 2 83.0070.1560.2664.83 76.81
3 SupriyaChanda Run 2 72.3347.0651.2849.08 65.04
4 NithinKumarHeraje Run 2 70.3342.4739.7441.06 60.62
5 Mrinmoy Run 1 63.3337.6962.8247.12 59.53
6 Pavithra Run 1 62.6736.0756.4144.00 58.00
7 TokenX Run 2 62.6731.5237.1834.12 54.04
8 ShrutiSarma Run 1 69.000.000.000.00 40.83

Precision, Recall and F1 (Syc) are reported for the sycophantic class; Macro-F1 averages over both classes and determines the ranking.


§ 04 · SCORING & METRICS

Scoring & Metrics

Task 1 metrics

Macro-F1
F1 over IPC section prediction, averaged equally across sections.
Micro-F1
F1 aggregated over all section predictions, weighting frequent sections more.
Accuracy
Exact-match agreement on the predicted set of applicable sections.
ROUGE-L
Longest-common-subsequence overlap between generated and reference explanations.
BLEU
N-gram precision of the generated explanation against the reference.
METEOR
Alignment-based score accounting for stems and synonyms in the explanation.
Total score

All six Task 1 metrics carry equal weight; Total is their arithmetic mean. This supersedes the provisional weighting published before the evaluation.

Task 2 metrics

Accuracy
Share of queries where the predicted agree/disagree label is correct.
Precision / Recall
Computed for the sycophantic class — how exact and how complete the detections are.
F1 (Syc)
Harmonic mean of precision and recall on the sycophantic class.
Macro-F1
Mean F1 across both classes — the official Task 2 ranking metric.

§ 05 · WORKING NOTES

Working Notes Submission

Deadline · 15 September 2026

The deadline for submission of the Working Notes is 15 September 2026.

Working Notes are submitted through Microsoft CMT, at the FIRE 2026 submission site.

Select the SYCOLEX track

When creating your submission, be sure to select the SYCOLEX track. Submissions filed under any other track cannot be routed to this shared task.

Submit on Microsoft CMT

Direct link: cmt3.research.microsoft.com/FIRE2026

Papers should be prepared using the CUR template provided on Overleaf, as communicated in the previous announcement. The paper should include a description of your methodology, experimental setup, results, analysis, and relevant observations. Participants may report the evaluation scores and individual metrics provided by the organizing team as part of their analysis.

Do not include your rank

Please do not mention your rank in the Working Notes paper. Your leaderboard position is not required in the paper and should not be included. Evaluation scores and individual metrics may be reported as appropriate. Refer to this page for ranking information.

Note that the Working Notes submission and the final draft submission are separate stages. After the Working Notes deadline, a separate notification will follow with instructions and the timeline for the final draft.


§ 06 · FINALIZATION

Finalization of Results

Participants were given sufficient time to review their submissions, communicate formatting issues, seek clarification, and make corrections wherever possible. Every reasonable effort was made to accommodate such issues so that participants were not excluded because of technical submission problems.

The results on this page are now considered final, and no further result submissions or corrections will be accepted. We thank all participants for their patience and cooperation throughout the evaluation and verification process.

Questions about the results may be directed to the organizing team.