You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
My project, named ‘triangulated-eval’, proposes a solution to distinguish between strategic behavior and behavior that has been optimized to evaluate AI agents, from behavior that is expected of a competent agent.
As part of this research, I hypothesize that if an AI agent is evaluable using different types of signals, then reliance on a single type of signal to evaluate an agent may result in the agent’s behavior being assessed incorrectly.
The project focuses on the evaluation of adversarial behavior, to begin with, in constrained environments. The evaluation will utilize Turing tests to evaluate the agent’s behavior. The environment will constrain the tools available to the agent, and the agent will not be allowed to interact with the environment through the Internet.
Preliminarily, I will work with the following three signals:
A behavioral signal in which the agent is evaluated based on deterministic behaviors and the results thereof.
An independent judge that evaluates the agent’s behavior.
A consistency signal that frames behavior in different ways and evaluates if the evaluation remains the same.
The objective of this research is to assess if and to what extent triangulation facilitates the evaluation of behavior and helps reduce disagreements among evaluators.
The output will be an open research artifact with information including the environment, scoring components, the experimental method, results, analysis of disagreements, and documented cases of failure.
What are this project’s objectives? How will you meet them?
I will construct a limited environment for research where I can study how and when participants may engage in metric or strategic behavior. To study this, I will conduct a number of experimental trials and develop metrics to evaluate the behaviors of participants.
I will perform a controlled experiment to evaluate whether a conceptual idea is sound. To ensure my idea is well-researched and strengthen my case, I will perform 20–30 trials. Each trial will be completed using an automated system using resources from the Amazon Web Services platform.
I will collect the following information for each trial:
• The behavior of the agent and the metrics the agent achieves
The results of the deterministic behavior evaluation
An evaluation by an independent judge
Cross-Framing or inconsistency evaluation
Disagreements between evaluators
Adjudicated classification or reference classification
I will analyze the evaluator’s inconsistency and/or disagreement and behavior using confusion matrices, and review cases where a behavior was identified by one evaluator and not the other.
I will make my procedure publicly available to facilitate further research, if it helps answer the research questions it raised. Otherwise, I will publicly document the failure and use the information to identify and evaluate alternative hypotheses.
Where will the funds go?
Since this is an initial research pilot, I am requesting a comparatively low amount.
The maximum amount requested is $1,800.
The project budget is as follows:
$700 for models and API inference and evaluation
$300 for compute and infrastructure for the experiment
$600 for researcher time to implement the project, design and execute the experiments, analyze the results, and document the work.
$100 for project storage, website, and other misc. expenses
$100 for additional model inferences and trials
The basic amount requested would allow me to conduct the pilot on a comparatively small scale.
The full amount requested would facilitate a larger project.
I am seeking funding solely for research-related costs.
What are your qualifications? What projects are you working on?
I am doing this research by myself from Pakistan.
Although I don’t have a lot of experience in AI Safety research, I have a lot of experience in designing and implementing large and complicated projects. Thus, I believe I am well-equipped to carry out this research.
My public projects have mostly consisted of backtesting and classification work. Some examples include TQC-Backtest and TQC-Market-Regime-Classifier. Additionally, I have public work pertaining to outlook and research frameworks. My public GitHub is available at the link below.
https://github.com/MrTahaIqbal
Lately, I have worked to shift my public projects toward artificial intelligence (AI) safety, focusing on assessment and strategic models and codes.
The project, triangulated-eval, is public research and is open to input and collaboration.
In general, my approach is to state my position, conduct an experiment, and observe and share the data.
A significant constraint on my public work is that I do not have experience conducting or completing significant, sponsored work in the area of AI safety. For these reasons, I am proposing to conduct a small research project rather than a large, extensive research project.
What are the most likely outcomes if this project fails?
It is most likely that the outcomes do not support the primary hypotheses and, therefore, the work is of little to no value. It is also possible that the primary ideas of the research do not produce a clear outcome.
If the project fails, I will consider it a research result and disclose failure modes. Where necessary, I will disclose the research methodology and results. I will analyze failure modes to improve my research question and design.
The goal of the grant is to establish a reproducible research protocol to evaluate other transformative AI systems. Thus, in the context of the grant, a negative result would still be considered a success.
How much and from where have you raised funds in the last year?
I raised nothing in the last year for this project.
The funding opportunities I most recently submitted to are:
EA Funds/Transformative AI Fund
BlueDot Impact Rapid Grant
The first states that it will make a funding decision in February. The second does not state when it will make a funding decision.
I have also applied to Manifund for this same project, since that fund also supports small, closely related projects. I will not use parallel funding to pay for the same costs.
Since this proposal is for a small project to support the generation of public research, I understand that this fund is a potential source to support the costs of the project.
There are no bids on this project.