MUMBAI: When it comes to artificial intelligence, sometimes knowing what you don’t know is the smartest move. That is precisely the idea behind a new benchmark developed by researchers at Eros GenAI, which has earned top honours at the Google DeepMind × Kaggle “Measuring Progress Toward AGI: Cognitive Abilities” Hackathon.
Eros Innovation announced that Eros GenAI researchers Arjun Thilak R. and Ramkumar M. V. were named among the four Grand Prize winners for creating GAUGE (Gap Between What LLMs Know and What They Do About It), a benchmark designed to evaluate whether large language models not only recognise uncertainty but also respond responsibly by refraining from answering when confidence is low.
Unlike conventional AI benchmarks that primarily measure accuracy, GAUGE examines whether a model’s behaviour aligns with its own confidence. The framework aims to address a critical challenge in deploying AI systems across sectors such as healthcare, education, government and enterprise, where confidently delivering an incorrect answer can have significant consequences.
The benchmark measures three key capabilities: how accurately a model evaluates its own confidence, whether it changes its behaviour when uncertain, and whether confidence and decision-making remain meaningfully connected.
To validate the framework, the researchers evaluated eight leading AI models from four model families using a three-stage metacognitive testing protocol covering mathematics, logical reasoning and factual knowledge.
Eros Innovation co-founder and co-president Ridhima Lulla said, “This recognition is an important milestone for Eros GenAI and demonstrates the calibre of original research emerging from our team. As AI systems become more powerful and increasingly autonomous, trust will depend not only on what a model knows, but on whether it understands the limits of that knowledge and behaves responsibly when uncertain. GAUGE directly supports our vision of building sovereign, culturally intelligent and accountable AI systems in which human oversight, safety and responsible action are designed into the architecture from the beginning.”
The company said it plans to incorporate the GAUGE methodology into its broader AI evaluation and assurance framework to assess global, open-weight and internally developed models across parameters such as reliability, uncertainty, safety and human escalation behaviour.
It also intends to expand the research into areas including Cultural Intelligence, rights-aware AI, education, citizen services and sovereign AI applications.
Eros GenAI ai research scientist Arjun Thilak R. said, “A model that knows it is uncertain but continues to act with confidence creates a serious risk for human operators. GAUGE was designed to measure the relationship between self-awareness and responsible action, rather than evaluating accuracy alone.”
Eros GenAI ai research scientist Ramkumar M. V. added, “GAUGE moves beyond conventional single-answer testing by examining how a model assesses uncertainty across multiple turns and whether that self-assessment changes its behaviour. This is especially important for real-world AI systems, where a reliable model must know when to proceed, when to abstain and when to seek human oversight.”
The recognition strengthens Eros Innovation’s broader ambition to build a sovereign Cultural Intelligence platform that combines AI orchestration, culturally aware models, rights and identity infrastructure, trustworthy AI agents and country-specific intelligence systems. With responsible AI becoming a defining priority for governments and enterprises alike, the company is betting that the future of artificial intelligence will be measured not only by how much it knows, but by how wisely it chooses to respond.