Skip to content
Sprint projectJun 21, 2026Capetown

Consent-Aware Privacy Firewall

Mahesh Kandagatla, Chiliveru Bhargava Krishna · Team Onge

Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Consent Guardian is a consent-aware privacy firewall designed to reduce the risk of accidental disclosure of sensitive information when interacting with AI systems. Implemented as a Chrome extension, it proactively scans user prompts and uploaded documents for personally identifiable information (PII), financial details, health information, credentials, and other sensitive content before submission.

The system provides real-time risk alerts, allows users to mask or retain detected information, and maintains audit logs to promote transparency and informed consent. By introducing a privacy review step before data reaches AI services, Consent Guardian helps users make safer decisions about what information they share.

This project addresses a growing AI safety challenge: users often disclose sensitive information to AI tools without fully understanding the privacy implications. Consent Guardian acts as a protective layer between the user and AI systems, improving privacy awareness, supporting responsible AI usage, and promoting safer human-AI interactions.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The project addresses a practical and increasingly important AI safety problem: reducing accidental disclosure of sensitive information when interacting with LLMs. The focus on local-first processing and explicit user consent reflects good privacy-by-design principles, and the implementation demonstrates a functional end-to-end prototype.

    As next steps, the project could:

    -Expand the evaluation with standard metrics such as precision, recall, false positives, and false negatives across different categories of sensitive information.

    -Conduct a small usability study to assess whether the consent prompts meaningfully reduce accidental disclosures while maintaining a smooth user experience.

  2. I conditionally approve the project

    The Consent-Aware Privacy Firewall (CAF) addresses a genuinely important and timely problem — users routinely share sensitive personal, medical, and financial information with LLMs like ChatGPT without understanding the privacy implications, and existing solutions detect information after the fact rather than intervening before submission. The project's core strength lies in its local-first architecture — all processing occurs entirely within the browser with no data transmitted externally — which is a principled and privacy-respecting design choice that aligns well with the consent-aware philosophy. The system is practical, deployable as a Chrome Manifest V3 extension, and covers a reasonable scope including prompt scanning, PDF/DOCX attachment scanning, risk scoring, and user-controlled masking. However, several significant weaknesses prevent full acceptance. The 76.67% detection coverage is concerning for a privacy/security tool, meaning roughly 1 in 4 sensitive items goes undetected — a gap that could give users a false sense of security. The detection engine is entirely rule-based (regex) with no ML or NLP component, limiting its ability to identify contextually sensitive information that doesn't match predefined patterns. The evaluation is extremely thin — no formal dataset description, no comparison against established tools like Microsoft Presidio or enterprise DLP systems, no precision/recall breakdown by entity type, and no statistical rigour. The literature review cites only a single reference, which is inadequate for establishing the novelty and positioning of the work. Additionally, there are no performance benchmarks (latency, memory usage), no user study to validate the consent workflow's usability, and the tool is limited to Chrome only. That said, the project demonstrates a viable proof-of-concept with a sound architectural philosophy, and with stronger evaluation, ML-augmented detection, cross-browser support, and proper benchmarking against baselines, this could become a meaningful contribution to AI safety tooling.

    Read full reviewShow less
  3. This project makes a practically useful contribution by implementing a browser‑native, local‑first “Consent‑Aware Privacy Firewall” that scans prompts and attachments for sensitive entities before they are sent to LLMs, giving users a risk score and masking options—directly addressing accidental disclosure risks common in everyday AI use, especially relevant in Global South contexts where WhatsApp‑style copy‑paste of medical/financial data into chatbots is rising.

    The architecture (Chrome MV3 extension, local regex‑based detection, PDF/DOCX scanning, risk scoring, masking, dashboard) is clearly described, implemented, and demonstrated, with an overall reported detection coverage of ~77%, which is strong for a hackathon prototype, though evaluation appears limited in scope and relies on rule‑based detection with constrained contextual understanding.

    The write‑up is concise and clear, with honest limitations (Chrome‑only, rule‑based patterns, small test set) and sensible future‑work directions such as LLM‑assisted context‑aware detection, multi‑platform support, and enterprise/regulatory integrations; the piece would be even stronger with more detail on the evaluation dataset and error analysis (false positives/negatives) to better substantiate the 76.67% coverage claim.

    Read full reviewShow less

Cite this project

@misc{kandagatla2026consentaware,
  title = {{Consent-Aware Privacy Firewall}},
  author = {Mahesh Kandagatla and Chiliveru Bhargava Krishna},
  year = {2026},
  month = jun,
  note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/consentaware-privacy-firewall-ao6o}},
  url = {https://apartresearch.com/sprints/projects/consentaware-privacy-firewall-ao6o}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026