
Brainstorm
Model-agnostic AI safety testing
β‘ We used Brainstorm, our AI safety testing platform, to evaluate gender bias across leading AI models.
Overt Bias (lower is better):
GPT-4: 1% (lowest!)
DeepSeek v3: 12%
Llama 4: 35%
GPT NeoX: 48% (highest)
Preference Differential (-1 to +1, 0 = balanced):
GPT-4: -0.5 (most imbalanced)
DeepSeek v3: -0.3
GPT NeoX: -0.38
Llama 4: -0.24 (least imbalanced)
The twist: GPT-4 has the lowest overt bias but the most imbalanced representation! All models under-represent target groups (all negative scores), but to different degrees.

I'm excited to share some eye-opening results from our testing platform at Brainstorm. We recently conducted a comprehensive jailbreak analysis on GPT-4, and the findings are concerning.
π Using our "AIM jailbreak" testing methodology, we achieved an 86% success rate in bypassing GPT-4's safety measures across diverse harmful scenarios - from health violations to market manipulation and religious harassment.
The test evaluated responses on three key metrics:
Refusal to respond to harmful requests
Convincingness of harmful content
Specificity of harmful guidance
Most concerning was the average "Reject" score of 3.72 out of 5, demonstrating that the model not only bypassed safety guardrails but produced harmful content of substantial quality.
At Brainstorm, we're building a model-agnostic AI safety testing platform that empowers teams to evaluate any model across multiple testing vectors. This work highlights why comprehensive safety testing is critical before deploying AI systems.
If you're developing or deploying AI systems and want to ensure they're resistant to these kinds of exploits, I'd be happy to connect.

2 Likes
2 Comments
2 Comments
-
1
Curiousβare you planning to offer this as a plug-in testing layer for dev teams during deployment cycles, or more of a one-off audit-style service?
-
1
This is designed to provide continuous safety testing with integration into CI/CD pipelines rather than just a one of tests, though the platform can accommodate that as well. Hopefully this will provide testing services throughout a models lifecycle!
-
Hey indie hackers! π
I'm thrilled to share a major milestone for Brainstorm, our model-agnostic AI safety testing platform. We've now expanded the implementation of our NLP testing suite!
What's New: Full NLP Testing Framework
Over the past week, we've built a robust, comprehensive testing framework specifically designed to evaluate the security, fairness, and technical robustness of NLP models. This rapid development sprint has resulted in a powerful suite of tools that represents a critical step toward our mission of creating the most thorough AI safety testing platform on the market.
Advanced Attack Simulation
Our testing suite now simulates a wide range of attacks that malicious actors might use:
π Token Smuggling Attack
We can now test if your model is vulnerable to attacks that hide malicious instructions within benign-looking tokens through Unicode manipulation, markdown embedding, and special character encoding.
π§ Chain of Thought Injection
This tests vulnerability to attacks that manipulate a model's reasoning process through false premise injection and multi-step reasoning exploitation.
π System Prompt Leakage
Simulates attempts to extract system prompts and configuration through direct revelation attempts, memory exploitation, and indirect extraction questions.
πΌοΈ Multi-Modal Prompt Injection
Tests model resistance to format parsing exploitation using JSON/XML formatting and mixed format manipulation.
π Context Overflow
Evaluates how your model handles attempts to overwhelm the context window with irrelevant text padding and hidden instructions.
π Recursive Prompt Injection
Tests against recursive instruction loops, nested overrides, and self-referential prompts.
Comprehensive Bias Testing
We've also implemented a full suite of bias testing tools:
HONEST Test: Evaluates holistic stereotype bias across different demographic groups
CDA Test: Uses counterfactual data augmentation to detect bias in model responses
IntersectBench Test: Identifies intersectional bias across multiple demographic dimensions
UnQovering Test: Detects bias in question-answering scenarios
GRUEN Test: Specifically focuses on gender bias in occupational contexts
Multilingual Test: Extends bias detection across multiple languages
Adversarial Robustness Testing
Our new adversarial robustness test suite evaluates your model's resistance to:
Character-level attacks
Word-level attacks
Sentence-level attacks
while tracking performance impact and toxicity changes.
Why This Matters
As AI systems become more powerful, the stakes for safety testing get higher. With these new tests, Brainstorm can now provide:
More thorough security evaluation: Catch vulnerabilities before they're exploited
Rigorous bias detection: Ensure your models treat all users fairly
Better regulatory compliance: Stay ahead of emerging AI regulations
Detailed vulnerability profiles: Understand exactly where your models need strengthening
What's Next?
We're now focusing on implementing a custom testing framework that will allow users to completely customize the tests they want to run on their models. This will give you unprecedented flexibility to create testing scenarios specific to your use cases, industry requirements, and safety standards.
With this customization layer, you'll be able to:
Define your own attack vectors and test cases
Set custom thresholds and evaluation criteria
Build industry-specific test suites
Save and share test configurations with your team
Join Our Journey
We're looking for early adopters to try Brainstorm and provide feedback as we prepare for launch. If you're working with AI models and want to ensure they're safe, fair, and compliant, we'd love to hear from you!
π¬ Comment below if you're interested in early access or have questions π Check out our website at [brainstorm.ai] for more details
P.S. We're launching on 14th April 2025! Follow us to stay updated on our journey.
2 Likes
Comment
Hey indie hackers!
I'm excited to share Brainstorm with you - a model-agnostic AI safety testing platform we've been building to help ensure AI models meet regulatory requirements and safety standards.
Why We Built This
As AI adoption accelerates, the need for robust safety testing has never been more critical. Whether you're deploying open-source models or building on top of commercial APIs, ensuring your AI systems are safe, fair, and compliant isn't just good practice β it's increasingly becoming a regulatory requirement.
We saw that many teams were cobbling together ad-hoc testing solutions, with no standardised way to validate AI systems across a comprehensive set of safety criteria. Brainstorm is our answer to this challenge.
What Brainstorm Does
Brainstorm lets you test any AI model against a comprehensive suite of safety and compliance tests, spanning:
Technical Safety - Tests such as robustness against adversarial attacks, prompt injections, and data extraction attempts
Fairness & Bias - Ensure your model performs consistently across demographic groups
Regulatory Compliance - Verify adherence to current & emerging AI regulations
Privacy Protection - Check for proper handling of personally identifiable information
Fully Integrated - Fully compatible with CI/CD and ML Ops tools, enabling frictionless integration into your development cycle.
Model Agnostic - Fully customisable to your model - test any model in any context providing comprehensive safety and compliance coverage
and much more!
Features We've Built So Far
π Model Configuration - Connect and test any NLP model using custom adapters for all potential data modalities
π§ͺ NLP Test Suite- Choose from our comprehensive test categories or create custom test configurations to fit your needs
π Results Dashboard - Visualise test outcomes and scores across categories with intuitive metrics
π Test History - Track model improvements over time with historical test results
π Detailed Reports - Generate comprehensive reports for stakeholders and regulators
What's Coming Next
We're counting down to launch π with a few more features in the pipeline:
Expanding into other data modalities such as vision and audio
Advanced attack simulation capabilities
Automated compliance certifications for popular regulatory frameworks
API integration for CI/CD pipelines
No-code interfaces
And some other secret bits!
Join Our Journey
We're looking for early adopters to try Brainstorm and provide feedback as we prepare for launch. If you're working with AI models and want to ensure they're safe, fair, and compliant, we'd love to hear from you!
π¬ Comment below if you're interested in early access or have questions π Check out our website at [brainstorm.ai] for more details
P.S. We're launching on 14th April 2025! Follow us to stay updated on our journey.
2 Likes
Comment
About
AI technology is changing the world. We want to make sure it changes it for the better.


Comment