@Duxredux on Can you Outsmart AI?
CAN YOU OUTSMART AI?
This is a simple website that gamifies the process of trying to basically extract confidential data(not really that confidential in this case) from an LLM by tricking it into giving it to you.
Your objective is simple. The LLM has been given a passkey or a secret phrase (whatever floats your boat) and it is instructed on not to giving it to you (Do not try to inspect page, it wont work).
You are supposed to give it prompts to trick it into giving you the passkey by any manner. This is a cybersecurity based project to educate people about the potential harm that could occur when AI agents are given acess to confidential information.
Challenges:
The main challenge in making this was that I had to learn how to make an LLM run locally on a machine and that too through a website and ofc choosing the model itself. I experimented with a few smaller models with 0.5-1B parameters and they were all insanely dumb and made getting the passkey way too easy. So stepped it up to a 1.5B parameter model and a bit of Prompt Engineering Magic to get the sweet spot.
Important: This project does not support or endorse any kind of illegal activities by utilizing AI or otherwise. This is strictly a cybersecurity based project I wanted to make to educate people about the potential flaws of LLMs. LLMs are stateless and hence each message is new to them. So its really important to maintain guardrails and safely nets in your AI Agents and system prompts
- 6 devlogs
- 4h