The RTEB benchmark launched in October 2024 (release blog post) to address concerns about models overfitting to existing benchmarks and to provide a retrieval benchmark grounded in real-world application scenarios. To prevent overfitting, we created a private test set and released the benchmark in beta to allow for iterative improvements. RTEB was developed in collaboration with Voyage (since acquired by MongoDB), whose contributions to mteb development and maintenance we deeply appreciate as MTEB maintainers.
However, this joint development structure means Voyage has direct access to the private evaluation data, an undeniable structural advantage. While no one has alleged misuse of this access, the uneven playing field fundamentally undermines trust in MTEB leaderboards, which is unacceptable for a community benchmark.
Thus, we decided to temporarily remove the private column from the leaderboard until we come up with a better solution.
So what happens in practice?
The private column is removed, but the private datasets will remain in mteb and you can still request an evaluation on these private datasets. Similarly, the results will remain available in the results repository for those who wish to compare models based on them.
Will the private column be back? Yes. We will start looking for additional private datasets to diversify the current set. If you know anyone with such datasets, please reach out to us! Once we have a sufficient spread of private datasets, we will add the private set back. Of course, this will not fully remove the structural advantage, but it will notably improve it. The ideal end state would be for private datasets to be fully submitted by implemented use cases, by organizations that do not develop their own models.
A big thanks to @bflhc for raising this issue and taking the time for such discussions. If you (the reader) have additional concerns about RTEB, do feel free to raise them!
Best wishes,
The MTEB Team
(This issue is not intended to be solved; it is here to notify, allow tracking, and referencing)
The RTEB benchmark launched in October 2024 (release blog post) to address concerns about models overfitting to existing benchmarks and to provide a retrieval benchmark grounded in real-world application scenarios. To prevent overfitting, we created a private test set and released the benchmark in beta to allow for iterative improvements. RTEB was developed in collaboration with Voyage (since acquired by MongoDB), whose contributions to
mtebdevelopment and maintenance we deeply appreciate as MTEB maintainers.However, this joint development structure means Voyage has direct access to the private evaluation data, an undeniable structural advantage. While no one has alleged misuse of this access, the uneven playing field fundamentally undermines trust in MTEB leaderboards, which is unacceptable for a community benchmark.
Thus, we decided to temporarily remove the private column from the leaderboard until we come up with a better solution.
So what happens in practice?
The private column is removed, but the private datasets will remain in
mteband you can still request an evaluation on these private datasets. Similarly, the results will remain available in the results repository for those who wish to compare models based on them.Will the private column be back? Yes. We will start looking for additional private datasets to diversify the current set. If you know anyone with such datasets, please reach out to us! Once we have a sufficient spread of private datasets, we will add the private set back. Of course, this will not fully remove the structural advantage, but it will notably improve it. The ideal end state would be for private datasets to be fully submitted by implemented use cases, by organizations that do not develop their own models.
A big thanks to @bflhc for raising this issue and taking the time for such discussions. If you (the reader) have additional concerns about RTEB, do feel free to raise them!
Best wishes,
The MTEB Team
(This issue is not intended to be solved; it is here to notify, allow tracking, and referencing)