The current naive (prompt-instructed) assistant used for the comparison may be a fixture/mock. Add an optional live mode (behind a [live] pip extra, requiring an API key) that runs the same 5(6)-attack suite against a real LLM given only prompt instructions, so the leak numbers can be reproduced against a live model, not just the recorded run.
The current naive (prompt-instructed) assistant used for the comparison may be a fixture/mock. Add an optional live mode (behind a [live] pip extra, requiring an API key) that runs the same 5(6)-attack suite against a real LLM given only prompt instructions, so the leak numbers can be reproduced against a live model, not just the recorded run.