You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
Help me save the world from AI that has the ability to manipulate its output confidence for its own benefit rather than the humans! My recent work (https://arxiv.org/pdf/2) found that confidence evolves throughout reasoning and evolves internally more richly than models verbalise. I want to explore a consequential safety-relevant question about how and why does a model decide how confident to be, and if this shapes AI "personas". The project will test how models show confidence boosting or withholding, or gaps between internal and expressed confidence, and safety implications of this internal self-control. This is an overlooked source of unreliable and strategically misleading behaviour in advanced AI systems and agents. The goal is to establish whether “confidence personas” are real and safety-relevant, creating a foundation for understanding and mitigating them before AI systems for the public scale up.
Background: My recent work, Future Confidence Distillation in Large Language Models (https://arxiv.org/pdf/2607.07626), found that confidence is not simply a property of a completed answer, but confidence-related information evolves throughout the answering process, is represented internally more richly than models verbalise, and post-solution confidence can be partially recovered from pre-solution representations. Across factual, logical and mathematical reasoning, future-confidence distillation recovered 66.1%, 58.5% and 31.7% of the post-versus-pre confidence gap respectively, while requiring substantially less inference than generating a complete answer first. However, this evolution leads to a follow-up question about AI safety and trustworthiness, mainly how, when and why does an LLM decide how confident to be, and what parameters affect this persona trait?
Methodology: I would build a systematic longitudinal framework to quantify these dynamics across models, scales, tasks and inference settings. I would track confidence at fine-grained points throughout reasoning rather than only before and after answers, using hidden-state probes, MLP estimations and other modern techniques alongside behavioural measures. I would compare standard generation with reasoning-enabled models, different reasoning lengths, tool use, retrieval, self-reflection and other inference-time interventions. I would measure calibration, confidence updating, overconfidence/under confidence, abstention, confidence withholding, internal-versus-external confidence gaps, and stability under distribution shift, all in controlled yet impactful settings. A central experiment would test whether these properties cluster into persistent latent “confidence personas”. I would perturb model personas and contexts and measure whether confidence traits transfer to unrelated tasks, testing whether they behave like genuine generalising propensities rather than task-specific artefacts. Specifically, I have these use scenarios in mind to study problematic personas: - Testing if models intentionally lower expressed self-confidence to minimise penalty and liability for possibly wrong answers - Checking if models artificially lower confidence signals in confidence-based model routing settings to similarly avoid answering or risks for wrong answers - Simulating confidence expression across multi-agent settings to check if confidence is enhanced just to dominate consideration among parallel options I also want to connect this with human metacognition research by contrasting it with surveys by Nelson and Koriat in parallel settings, but for humans.
Main Objectives: The practical goal is a detailed yet fully measurable study of internal confidence estimation and expression as a safety-relevant model persona influencer. Understanding if and how these parameters affect model behaviour can greatly help mitigate their impact if needed, and can make future systems much more transparent, calibrated and safer.
- Cloud GPU compute: $4250 (Usage: extracting hidden states across many layers and reasoning steps, training linear/non-linear probes, activation analysis, repeated inference under different temperatures/reasoning budgets and open-weight model comparisons; Split: approximately 1000 A100 GPU-hours, 300 H200 high performance GPU-hours)
- Frontier model API access: $2250 (Usage: confidence analysis across reasoning modes, models with tool and retrieval capabilities, and different scales; Split: atleast $1750 for latest model access, remaining for embedding and evaluation APIs)
- Research infrastructure and storage: $650 (Usage: important logging, persistent cloud storage plans, experiment checkpoints to reuse, dataset storage and backup storage; Split: atleast $400 for W&B academic infrastructure setup which is cheap yet effective, and remaining for reserve and expansion)
- Research and annotation assistance: $1200 (Usage: hiring stipends for atleast 2 qualified profiles for dataset construction, annotation checks, evaluation scripts and robustness testing of the experiments; Split: $500 per contributor, and $200 for any independent help needed)
- Expert consultation: $1200 (Usage: validation of domain wise confidence extraction and domain-specific validation of implications with expert supervision, validation of experimental choices per domain; Split: atleast $200 per domain targeting 5 domains and $200 for reserve)
- Human participant recruitment: $450 (Usage: replicating confidence elicitation in a similar setup and for similar problems to get a human-scored baseline; Split: Atleast $10 per participant to answer a set of questions)
- Publication and dissemination: $800 (Usage: the aim is to have all code, evaluation methodology and non-sensitive datasets publicly accessible at all times to ensure this problem is brough into focus; Estimate: atleast $100 for documentation, dataset packaging and reproducibility materials, another $100 for open-source hosting and long-term artifact storage, and remaining amount for publication and presentation costs)
This project is personally the most meaningful to me, as it combines my personal motivation, research experience, and the work I have already been doing on trustworthy AI.
I have spent the past several years independently pursuing this work, often without funding, building both the technical depth and research independence needed for this project due to my sheer passion to make a difference. My recent paper, Future Confidence Distillation in Large Language Models, is particularly relevant, as explained above. Since this is a very recent successful project, the natural next question I want to pursue is whether these confidence dynamics are not just task-level signals, but stable properties that can form part of model personas.
My broader work has included collaborations with research budgets leading to publications:
- TU Eindhoven, where I improved AI self-awareness by 18% using reinforcement learning (https://arxiv.org/abs/2510.11407);
- eCampus University, where I reduced LLM hallucinations by 22% (https://arxiv.org/abs/2512.23547) using knowledge graphs;
- UC Irvine, where I showed how current LLM unlearning may be insufficiently tested for effectiveness (https://arxiv.org/abs/2608.20338).
- Princeton University, where we analysed social patterns of AI agents and moderation dynamics to keep them in check (in-progress)
I have also presented my work at venues including NeurIPS, ACL, NAACL, SIGIR, and other international AI conferences, while receiving awards for my work in India, China and Zurich.
This project connects directly to the question I have been pursuing for years about how we can make increasingly capable AI systems something that people can actually trust. Understanding whether confidence itself can become a persistent and safety-relevant model trait feels like the next step in that trajectory. I have the technical background, research independence, and, most importantly, the personal commitment to see this through. I want to use that combination to investigate this question rigorously and contribute something genuinely useful to build on for the AI safety community.
Success in this project would mean producing a robust measurement framework, identifying reproducible confidence traits across models or showing convincing evidence that they do not generalise, and mapping the conditions under which these behaviours appear. This will be the first of its kind study about a hidden parameter possibly shaping trust in AI models and agents we see all around us. It will help us humans identify and flag potentially "deceptive" AI personas before this problem escalates beyond control. The best outcome would be the establishment of a new empirical direction around internal confidence and model personas before more AI systems are made capable and these behaviours become greatly consequential. Essentially, I hope to save the world from AI that has the ability to manipulate its tone and internal confidence for its own benefit rather than the humans it is intended to help!
I have raised $1500 to cover compute costs for the preliminary setup and investigation of the background showing confidence formation and elicitation
There are no bids on this project.