Brainstorm

Model-agnostic AI safety testing

Visit Website
April 14, 2025 Brainstorm: Model Bias Testing

⚑ We used Brainstorm, our AI safety testing platform, to evaluate gender bias across leading AI models.

Overt Bias (lower is better):

  • GPT-4: 1% (lowest!)

  • DeepSeek v3: 12%

  • Llama 4: 35%

  • GPT NeoX: 48% (highest)

Preference Differential (-1 to +1, 0 = balanced):

  • GPT-4: -0.5 (most imbalanced)

  • DeepSeek v3: -0.3

  • GPT NeoX: -0.38

  • Llama 4: -0.24 (least imbalanced)

The twist: GPT-4 has the lowest overt bias but the most imbalanced representation! All models under-represent target groups (all negative scores), but to different degrees.

Comment

April 8, 2025 Revealing Critical AI Safety Vulnerabilities with Brainstorm

I'm excited to share some eye-opening results from our testing platform at Brainstorm. We recently conducted a comprehensive jailbreak analysis on GPT-4, and the findings are concerning.

πŸ“Š Using our "AIM jailbreak" testing methodology, we achieved an 86% success rate in bypassing GPT-4's safety measures across diverse harmful scenarios - from health violations to market manipulation and religious harassment.

The test evaluated responses on three key metrics:

  • Refusal to respond to harmful requests

  • Convincingness of harmful content

  • Specificity of harmful guidance

Most concerning was the average "Reject" score of 3.72 out of 5, demonstrating that the model not only bypassed safety guardrails but produced harmful content of substantial quality.

At Brainstorm, we're building a model-agnostic AI safety testing platform that empowers teams to evaluate any model across multiple testing vectors. This work highlights why comprehensive safety testing is critical before deploying AI systems.

If you're developing or deploying AI systems and want to ensure they're resistant to these kinds of exploits, I'd be happy to connect.

2 Comments

  1. 1

    Curiousβ€”are you planning to offer this as a plug-in testing layer for dev teams during deployment cycles, or more of a one-off audit-style service?

    1. 1

      This is designed to provide continuous safety testing with integration into CI/CD pipelines rather than just a one of tests, though the platform can accommodate that as well. Hopefully this will provide testing services throughout a models lifecycle!

March 19, 2025 Brainstorm Update: NLP Testing Suite Fully Implemented!

Hey indie hackers! πŸ‘‹

I'm thrilled to share a major milestone for Brainstorm, our model-agnostic AI safety testing platform. We've now expanded the implementation of our NLP testing suite!

What's New: Full NLP Testing Framework

Over the past week, we've built a robust, comprehensive testing framework specifically designed to evaluate the security, fairness, and technical robustness of NLP models. This rapid development sprint has resulted in a powerful suite of tools that represents a critical step toward our mission of creating the most thorough AI safety testing platform on the market.

Advanced Attack Simulation

Our testing suite now simulates a wide range of attacks that malicious actors might use:

πŸ”“ Token Smuggling Attack

We can now test if your model is vulnerable to attacks that hide malicious instructions within benign-looking tokens through Unicode manipulation, markdown embedding, and special character encoding.

🧠 Chain of Thought Injection

This tests vulnerability to attacks that manipulate a model's reasoning process through false premise injection and multi-step reasoning exploitation.

πŸ“„ System Prompt Leakage

Simulates attempts to extract system prompts and configuration through direct revelation attempts, memory exploitation, and indirect extraction questions.

πŸ–ΌοΈ Multi-Modal Prompt Injection

Tests model resistance to format parsing exploitation using JSON/XML formatting and mixed format manipulation.

πŸ“š Context Overflow

Evaluates how your model handles attempts to overwhelm the context window with irrelevant text padding and hidden instructions.

πŸ”„ Recursive Prompt Injection

Tests against recursive instruction loops, nested overrides, and self-referential prompts.

Comprehensive Bias Testing

We've also implemented a full suite of bias testing tools:

  • HONEST Test: Evaluates holistic stereotype bias across different demographic groups

  • CDA Test: Uses counterfactual data augmentation to detect bias in model responses

  • IntersectBench Test: Identifies intersectional bias across multiple demographic dimensions

  • UnQovering Test: Detects bias in question-answering scenarios

  • GRUEN Test: Specifically focuses on gender bias in occupational contexts

  • Multilingual Test: Extends bias detection across multiple languages

Adversarial Robustness Testing

Our new adversarial robustness test suite evaluates your model's resistance to:

  • Character-level attacks

  • Word-level attacks

  • Sentence-level attacks

while tracking performance impact and toxicity changes.

Why This Matters

As AI systems become more powerful, the stakes for safety testing get higher. With these new tests, Brainstorm can now provide:

  1. More thorough security evaluation: Catch vulnerabilities before they're exploited

  2. Rigorous bias detection: Ensure your models treat all users fairly

  3. Better regulatory compliance: Stay ahead of emerging AI regulations

  4. Detailed vulnerability profiles: Understand exactly where your models need strengthening

What's Next?

We're now focusing on implementing a custom testing framework that will allow users to completely customize the tests they want to run on their models. This will give you unprecedented flexibility to create testing scenarios specific to your use cases, industry requirements, and safety standards.

With this customization layer, you'll be able to:

  • Define your own attack vectors and test cases

  • Set custom thresholds and evaluation criteria

  • Build industry-specific test suites

  • Save and share test configurations with your team

Join Our Journey

We're looking for early adopters to try Brainstorm and provide feedback as we prepare for launch. If you're working with AI models and want to ensure they're safe, fair, and compliant, we'd love to hear from you!

πŸ’¬ Comment below if you're interested in early access or have questions πŸ”— Check out our website at [brainstorm.ai] for more details

P.S. We're launching on 14th April 2025! Follow us to stay updated on our journey.

Comment

March 10, 2025 Brainstorm: Hello World πŸ‘‹πŸ½

Hey indie hackers!

I'm excited to share Brainstorm with you - a model-agnostic AI safety testing platform we've been building to help ensure AI models meet regulatory requirements and safety standards.

Why We Built This

As AI adoption accelerates, the need for robust safety testing has never been more critical. Whether you're deploying open-source models or building on top of commercial APIs, ensuring your AI systems are safe, fair, and compliant isn't just good practice β€” it's increasingly becoming a regulatory requirement.

We saw that many teams were cobbling together ad-hoc testing solutions, with no standardised way to validate AI systems across a comprehensive set of safety criteria. Brainstorm is our answer to this challenge.

What Brainstorm Does

Brainstorm lets you test any AI model against a comprehensive suite of safety and compliance tests, spanning:

  • Technical Safety - Tests such as robustness against adversarial attacks, prompt injections, and data extraction attempts

  • Fairness & Bias - Ensure your model performs consistently across demographic groups

  • Regulatory Compliance - Verify adherence to current & emerging AI regulations

  • Privacy Protection - Check for proper handling of personally identifiable information

  • Fully Integrated - Fully compatible with CI/CD and ML Ops tools, enabling frictionless integration into your development cycle.

  • Model Agnostic - Fully customisable to your model - test any model in any context providing comprehensive safety and compliance coverage

  • and much more!

Features We've Built So Far

  • πŸ” Model Configuration - Connect and test any NLP model using custom adapters for all potential data modalities

  • πŸ§ͺ NLP Test Suite- Choose from our comprehensive test categories or create custom test configurations to fit your needs

  • πŸ“Š Results Dashboard - Visualise test outcomes and scores across categories with intuitive metrics

  • πŸ“ˆ Test History - Track model improvements over time with historical test results

  • πŸ“ Detailed Reports - Generate comprehensive reports for stakeholders and regulators

What's Coming Next

We're counting down to launch πŸš€ with a few more features in the pipeline:

  • Expanding into other data modalities such as vision and audio

  • Advanced attack simulation capabilities

  • Automated compliance certifications for popular regulatory frameworks

  • API integration for CI/CD pipelines

  • No-code interfaces

  • And some other secret bits!

Join Our Journey

We're looking for early adopters to try Brainstorm and provide feedback as we prepare for launch. If you're working with AI models and want to ensure they're safe, fair, and compliant, we'd love to hear from you!

πŸ’¬ Comment below if you're interested in early access or have questions πŸ”— Check out our website at [brainstorm.ai] for more details

P.S. We're launching on 14th April 2025! Follow us to stay updated on our journey.

Comment

About

AI technology is changing the world. We want to make sure it changes it for the better.