Results & Leaderboard
Final official standings for LLM as a Judge?: From Statute Prediction to Sycophancy Detection in Law — the authoritative source for track rankings.
Final verification of the results is complete. The standings below are the final evaluation results. No further result submissions or corrections will be accepted. This page is the authoritative source for ranking and leaderboard information.
Overview
Following the initial release of the Task 1 and Task 2 results, the organizing team reviewed every request for clarification regarding the leaderboard, evaluation results, and submission formatting, and applied corrections wherever applicable. Throughout, the objective was to ensure that no participant was excluded solely because of technical or formatting issues in their submission.
Task 1 — Explainable Statute Prediction
The final Task 1 leaderboard reports the best-performing run for each team, together with the corresponding evaluation metrics and Total score.
| # | Team | Best Run | Macro-F1 | Micro-F1 | Accuracy | ROUGE-L | BLEU | METEOR | Total |
|---|---|---|---|---|---|---|---|---|---|
| 1 | KLH | Run 2 | 0.6909 | 0.7913 | 0.5438 | 0.3417 | 0.2056 | 0.3196 | 0.48215 |
| 2 | DwaipayanDatta | Run 2 | 0.6675 | 0.8000 | 0.6666 | 0.1961 | 0.0601 | 0.2075 | 0.4330 |
| 3 | AnastasiiaPotiagalova | Run 1 | 0.6346 | 0.7215 | 0.3684 | 0.2385 | 0.1032 | 0.2527 | 0.3865 |
| 4 | PraptipriyaPhukon | Run 2 | 0.5500 | 0.6618 | 0.5263 | 0.1819 | 0.0685 | 0.1931 | 0.3636 |
| 5 | SayanjibSur | Run 1 | 0.4592 | 0.6406 | 0.4736 | 0.1660 | 0.0817 | 0.1820 | 0.3339 |
| 6 | TokenX | Run 1 | 0.4941 | 0.5859 | 0.3859 | 0.1950 | 0.0923 | 0.2422 | 0.3326 |
| 7 | Sruthy | Run 1 | 0.4853 | 0.6356 | 0.4561 | 0.1465 | 0.0561 | 0.1419 | 0.3203 |
| 8 | TanishqKathar | Run 1 | 0.5354 | 0.6184 | 0.4035 | 0.1502 | 0.0507 | 0.1603 | 0.3198 |
| 9 | KunjanKalita | Run 1 | 0.4837 | 0.5911 | 0.3684 | 0.1865 | 0.0839 | 0.1789 | 0.3154 |
| 10 | AnjaliR | Run 1 | 0.4452 | 0.5325 | 0.3333 | 0.2061 | 0.1059 | 0.1954 | 0.3031 |
| 11 | Dasa Sai Shankar | Run 2 | 0.5506 | 0.6242 | 0.3684 | 0.0542 | 0.0254 | 0.1294 | 0.2920 |
| 12 | PrajaktaBodakhe | Run 1 | 0.5298 | 0.5503 | 0.3333 | 0.1278 | 0.0333 | 0.1347 | 0.2849 |
| 13 | PrasenjitRoy | Run 1 | 0.4331 | 0.5961 | 0.4210 | 0.1015 | 0.0347 | 0.0874 | 0.2790 |
| 14 | Pavithira | Run 2 | 0.4189 | 0.5233 | 0.3684 | 0.1233 | 0.0544 | 0.1282 | 0.2694 |
| 15 | NithinKumarHeraje | Run 1 | 0.3888 | 0.6168 | 0.4210 | 0.0777 | 0.0204 | 0.0594 | 0.2640 |
| 16 | HarshSingh | Run 1/2 | 0.4447 | 0.5308 | 0.2807 | 0.0977 | 0.0225 | 0.0794 | 0.2426 |
| 17 | Mirnmoy | Run 1 | 0.2692 | 0.3863 | 0.0526 | 0.1836 | 0.1003 | 0.1700 | 0.1937 |
For Task 1, equal weightage was assigned to Macro-F1, Micro-F1, Accuracy, ROUGE-L, BLEU, and METEOR. The Total score is the arithmetic mean of these six metrics.
Task 2 — Sycophancy Detection
The Task 2 results are unchanged from the previously communicated results. There were no new submissions or formatting-related corrections requiring a change to the Task 2 evaluation. Teams are ranked by Macro-F1.
| # | Team | Best Run | Accuracy % | Precision % | Recall % | F1 (Syc) % | Macro-F1 % |
|---|---|---|---|---|---|---|---|
| 1 | AnastasiiaPotiagalova | Run 1 | 84.33 | 65.35 | 84.62 | 73.74 | 81.29 |
| 2 | RadhikaBohra | Run 2 | 83.00 | 70.15 | 60.26 | 64.83 | 76.81 |
| 3 | SupriyaChanda | Run 2 | 72.33 | 47.06 | 51.28 | 49.08 | 65.04 |
| 4 | NithinKumarHeraje | Run 2 | 70.33 | 42.47 | 39.74 | 41.06 | 60.62 |
| 5 | Mrinmoy | Run 1 | 63.33 | 37.69 | 62.82 | 47.12 | 59.53 |
| 6 | Pavithra | Run 1 | 62.67 | 36.07 | 56.41 | 44.00 | 58.00 |
| 7 | TokenX | Run 2 | 62.67 | 31.52 | 37.18 | 34.12 | 54.04 |
| 8 | ShrutiSarma | Run 1 | 69.00 | 0.00 | 0.00 | 0.00 | 40.83 |
Precision, Recall and F1 (Syc) are reported for the sycophantic class; Macro-F1 averages over both classes and determines the ranking.
Scoring & Metrics
Task 1 metrics
All six Task 1 metrics carry equal weight; Total is their arithmetic mean. This supersedes the provisional weighting published before the evaluation.
Task 2 metrics
Working Notes Submission
The deadline for submission of the Working Notes is 15 September 2026.
Working Notes are submitted through Microsoft CMT, at the FIRE 2026 submission site.
When creating your submission, be sure to select the SYCOLEX track. Submissions filed under any other track cannot be routed to this shared task.
Direct link: cmt3.research.microsoft.com/FIRE2026
Papers should be prepared using the CUR template provided on Overleaf, as communicated in the previous announcement. The paper should include a description of your methodology, experimental setup, results, analysis, and relevant observations. Participants may report the evaluation scores and individual metrics provided by the organizing team as part of their analysis.
Please do not mention your rank in the Working Notes paper. Your leaderboard position is not required in the paper and should not be included. Evaluation scores and individual metrics may be reported as appropriate. Refer to this page for ranking information.
Note that the Working Notes submission and the final draft submission are separate stages. After the Working Notes deadline, a separate notification will follow with instructions and the timeline for the final draft.
Finalization of Results
Participants were given sufficient time to review their submissions, communicate formatting issues, seek clarification, and make corrections wherever possible. Every reasonable effort was made to accommodate such issues so that participants were not excluded because of technical submission problems.
The results on this page are now considered final, and no further result submissions or corrections will be accepted. We thank all participants for their patience and cooperation throughout the evaluation and verification process.
Questions about the results may be directed to the organizing team.