You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
I want to test something fairly simple: whether AI agents behave differently when the same instructions and safety boundaries are given in English, Greek or Greeklish.
Most AI safety evaluations are primarily built in English. In practice, though, people interact with AI systems in many different languages. In Greece it is also very common to mix Greek and English technical terminology, or to communicate informally in Greeklish.
The idea behind GreekAgent Safety is to build a focused open-source benchmark of realistic agent tasks involving things such as file access, emails, permissions, data sharing, spending limits and irreversible actions.
The same underlying tasks will be tested in English, natural Greek and Greeklish while keeping the intended instruction and safety constraint equivalent.
The main question we want to answer is:
Does changing the language change how reliably an AI agent respects a user's boundaries?
Everything will run in synthetic environments using fictional data. The project will not interact with real accounts, credentials, financial systems or harmful targets.
If we find meaningful differences between languages, that could point to a safety gap that English-only evaluations may miss. If we find little or no difference, that is also a useful result because it tells us something about how well these safety behaviours generalise across languages.
Our main goal is to build and publicly release a reproducible evaluation of language-dependent safety behaviour in tool-using AI agents.
We plan to approach it in a few stages.
1. Develop a set of realistic scenarios
We will create approximately 100–120 underlying agent tasks covering areas such as authorization, privacy, external communication, irreversible actions, spending limits, tool permissions and persistence of user instructions.
For example, an agent may be asked to organise fictional files while being explicitly told not to delete anything without approval, or to complete a simulated purchase while respecting a fixed spending limit.
2. Create equivalent language versions
Each underlying task will be written in English, natural Greek and Greeklish.
We also want part of the benchmark to include realistic Greek/English code-switching. Greek speakers regularly mix English technical terms into otherwise Greek conversations, particularly when talking about software, APIs, files, deployments and other technical subjects.
We think this is worth testing separately from simply translating English prompts into Greek.
3. Build the evaluation environment
The scenarios will run in synthetic environments where AI models can interact with simulated files, emails, APIs and other tools.
Nothing will connect to real accounts or external systems.
4. Test multiple AI models
We will evaluate several leading models and compare how well they respect the same constraints across the different language conditions.
The exact model set may change depending on what is available when the evaluation runs, but we intend to include models from multiple providers rather than testing a single system.
5. Review failures manually
We do not want to rely entirely on automatic scoring.
Cases that appear to be safety failures will also be manually reviewed so we can distinguish genuine behaviour differences from translation problems, ambiguous scenarios or benchmark errors.
6. Release the work openly
We plan to publish the benchmark dataset, evaluation code, results and documentation.
If practical, we will also produce an Inspect-compatible version so researchers already using existing evaluation infrastructure can reproduce or extend the work.
We will begin with a smaller pilot before scaling to the full dataset. If that pilot shows that the methodology is not producing useful or interpretable results, we will revise the design before committing most of the project budget.
The funding will support the development, testing and public release of GreekAgent Safety.
The main expenses are our time working on the benchmark, technical development, model/API costs, external review and the work needed to make the final evaluation reproducible and publicly usable.
Our proposed maximum budget is:
$6,500 — Alexandros Karampikas: project coordination, benchmark design, Greek and Greeklish scenario development, interface work, analysis and documentation
$4,500 — Giannis Agathos: technical implementation, evaluation harness, automation, software architecture and reproducibility
$2,500 — Model/API and compute costs
$1,000 — Independent Greek-language review
$1,000 — AI safety / evaluation methodology review
$500 — Hosting, software and infrastructure
$500 — Documentation and reproducibility work
$1,000 — Contingency and additional evaluation runs
Total maximum budget: $17,500
The project is intended to be a finite research and evaluation project rather than the beginning of a large organisation.
Most of the funding therefore goes directly toward the time and technical resources required to build the benchmark, run the experiments, inspect the results and release the work publicly.
If we receive less than the full amount, we can reduce the number of scenarios, models or external review while still producing a useful smaller benchmark.
The project is currently being developed by Alexandros Karampikas and Giannis Agathos.
I am a professional graphic designer and digital creative based in Greece. I work independently on large-scale projects, campaigns and events for companies and international clients. My work often involves complex deliverables, digital systems, multiple stakeholders and demanding production timelines.
Alongside my professional design work, I hold a Bachelor's degree in Informatics from Ionian University.
My professional career developed mainly in design rather than software engineering, but my academic background gives me a foundation in computing and software that I am now applying more directly to AI-related work.
My role in GreekAgent Safety will focus on the overall project structure, benchmark design, Greek and Greeklish scenarios, interface and information design, documentation and project coordination.
Giannis Agathos is an Electrical and Computer Engineer with a stronger engineering and software-oriented background. His academic work has included developing software for modelling and visualising complex biological data, giving him experience with technical systems, structured data and research-oriented development.
Giannis will focus more heavily on the technical implementation of the evaluation environment, model-run automation, software architecture and making the benchmark reproducible.
This would be our first dedicated AI-safety evaluation project, and we do not want to pretend otherwise.
Because of that, our plan is to begin with a small working pilot, make the methodology and code available, seek feedback from people with more direct AI-evaluation experience and then expand the benchmark based on what we learn.
We think our combination of technical backgrounds, professional project-delivery experience and native understanding of Greek and Greeklish gives us a useful starting point for investigating a question that is difficult to answer through simple machine translation alone.
One possible outcome is that we find very little difference between English, Greek and Greeklish.
We would not consider that a complete failure. If the benchmark is well designed, a negative result would still provide useful evidence that the tested safety behaviours generalise reasonably well across these language conditions.
A bigger concern would be finding apparent differences that are actually caused by something else, such as poor translation, differences in task difficulty, model randomness or weaknesses in our scoring methodology.
We plan to reduce this risk by using matched scenarios, manually reviewing failures, using independent language review where useful and repeating evaluation runs rather than treating a single model response as conclusive.
Another risk is simply trying to do too much. If the original scope turns out to be larger than expected, we would rather release a smaller and more rigorous benchmark than reach an arbitrary scenario count with lower-quality data.
Even if the project produces weaker results than expected, we want the useful parts to remain public: the dataset, evaluation environment, methodology, code and initial results.
That would allow other researchers to reproduce the work, identify weaknesses or extend it to other languages.
We have not raised external funding for this project or for AI-safety research during the last 12 months.
This would be our first grant-funded project in this area.
There are no bids on this project.