Skip to content
Sprint projectNov 24, 2025Ohio, SF, NYC

WikiGen: Bio‑logical safeguards for collaborative AI/ML on sensitive data

Wiktoria Leks · Team WikiGen

Submitted to Defensive Acceleration Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: WikiGen: Bio‑logical safeguards for collaborative AI/ML on sensitive data

Code (opens in new tab)More on wikigen.me (opens in new tab)
Share

Looking for data, model improvements, ie.. cures; WikiGen is an open sourced protocol connecting databases and a user’s data to allow selective consensus for a given inquiry, based on collective and/or private knowledge feeds. Particular to you, WikiGen evaluates a query, protocol or research objective, even an entire database submission against our internal base of researchers, commercial to academic partners and concerned healthcare consumers’ data repos. to securely and with privacy-first preserving means, discover missing components that exist already or possibilities for data collection campaigns as solutions (similar to bounties used in hacks). This way science, may not have to be repeated locally for the same vain, in vain, but rather shared to save time or money in R&D globally across. Animal model systems may not be needed to run experiments on, if the data exists already or if there is enough to simulate such animal models. Perhaps, the instrument you are about to optimize for a lab assay has been, before, performed exactly for your exact experiment constrains, however, not published or digitally accessible to your very credentials. Maybe you are a licensed MD or RN looking for insulin ASAP, or the rates of Asthma in children changing specific to your zip code. If the information is collected and exists in a database, all you need is a trustless way to request for it and prove you can be trusted! Currently, there is no way for researchers to “google” i.e. BLAST search align a SEQ file (GCAT to CGTA similarity) to locate plasmids or biological materials off of the input compliment target by percent homology of sequence (0-100% exact mach) overlap in data bases they have no authority to search in. Well yes.. for safety. However, in life science these databases are constantly being shared, in correctly logged, mailed and collaborated on. The models, benchmarks, scripts, standard operating procedure protocols and barcoded physical inventory carefully catalogued. How do we connect such very difficult to acquire, in the first place, tools and resources that may hold specific answers to someone else’s difficult question, if the data is collected and not shared, reused, or even not allowed to be open sourced. Such that If I work at BioMarin, I do not need to have friends in Amgen’s gene therapy group to know what sample plasmids they possibly may have in their freezer repertoire for their own IP campaigns, to be able to determine who to email for a sample. I'd much rather the void alert me of what I do not know, which I need based off of my current needs. Researchers and ML/AI agentic pipelines today require new methods to access data. Especially sensitive data that can be used for harm, without stunting the goodness that arrises from research aims, when utilized for the intended purposes. In antibody discovery, neuroscience research, oncology targets and in metabolomics, sequenced library sets are almost barely ever publicly added to open source gene banks. Publications in biology, such as Nature have priorly risked the identities of patients by publishing tissue sequence of sc-RNA matrices data, which other individuals were able to de-identify the patients. Publishing costs money and occasionally you get a very mean or confused reviewer, preventing critical relay channels to share important data or protocols in a heist matter. Lastly, bioinformatic or computational biology data scientists, unite in complaints of code not published to the paper, or not the right files or script versions publicly available. Institutions or biopharma corporations are only held accountable to incentivize cross dynamic synergy benefits to their entity's already self involved and financed collaborations, even when policy and laws exist such as to share failed clinical trial data (the U.S. states as law), they fail in implementation : clinicaltrials.gov does not support, nor has anyone significant noticed this is not being maintained, since 2013, for missing reports of clinical trial master files to be required for upload into an active registry for the, permission approved, collective to make use.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

Does this reduce AI-related catastrophic or existential risks?

Scoring guide
  1. 1Minimal Impact. The project has minimal relevance to AI safety. It doesn't meaningfully address risks uniquely enabled or accelerated by advanced AI systems.
  2. 2Tangential Connection. The project touches on AI safety concepts but lacks depth or specificity. The connection to AI-enabled threats (bio, cyber, or AI misuse) is weak or unclear.
  3. 3Clear AI Safety Value. The project clearly reduces AI-related risks with valuable contributions. It addresses specific threats from AI systems and engages meaningfully with biosecurity, cybersecurity, or AI safety challenges.
  4. 4Significant Impact Potential. The above, plus the project demonstrates scalable safety mechanisms or defensive approaches. It shows clear potential to buy time for solving harder problems like alignment, or creates positive externalities for the broader AI safety ecosystem.
  5. 5Major Advancement. The above, plus the project represents a significant leap forward in defensive AI safety. Judges would eagerly share this with biosecurity, cybersecurity, or AI safety researchers and expect it to influence the field.

Does this strengthen the shield against AI-enabled threats?

Scoring guide
  1. 1Minimal Relevance. The project is only tangentially related to defensive technology or societal protection. Connection to biosecurity, cybersecurity, or defensive infrastructure is unclear or missing.
  2. 2Some Relevance. The project has some relevance to defensive acceleration, but the connection is broad or generic. It touches on defense without specific focus on AI-enabled threats or protective capabilities.
  3. 3Clear Relevance. The project clearly addresses defensive gaps against AI-enabled threats. It connects to at least one track (biosecurity, cybersecurity, or defense infrastructure) and demonstrates understanding of the threat landscape.
  4. 4Strong Contribution. The above, plus the project builds on existing defensive approaches and offers novel tools, frameworks, or implementations. It explicitly explains how it strengthens defensive capabilities with realistic deployment potential.
  5. 5Breakthrough Impact. The above, plus the project provides breakthrough insights or tools that could significantly influence defensive technology development. It identifies critical gaps and presents compelling solutions with clear paths from prototype to deployed system.

Did you build something that actually works?

Scoring guide
  1. 1Incomplete or Flawed. The project appears rushed or incomplete. Technical implementation is flawed, core functionality doesn't work, or the approach is fundamentally unsound. Little to no documentation.
  2. 2Basic Competence. The project shows reasonable effort with basic technical competence. Core functionality partially works. Documentation exists but may be incomplete. Some limitations are acknowledged.
  3. 3Solid Hackathon Project. The project is technically solid and well-scoped for 48 hours. Core functionality works and is documented. Code/methods are understandable and limitations are honestly addressed. This is what a good weekend prototype should look like.
  4. 4Impressive Implementation. The above, plus the implementation exceeds typical hackathon quality. Clear methodology, thorough documentation, and working demo. The tool/prototype is immediately useful for defenders and could realistically be built upon.
  5. 5Exceptional Execution. The project far exceeds expectations with exceptional technical execution. The implementation is elegant, fully functional, and includes something special (e.g., deployed demo, exceptional documentation, innovative architecture, or clear startup potential).

No public critique yet.

Cite this project

@misc{leks2025wikigen,
  title = {{WikiGen: Bio‑logical safeguards for collaborative AI/ML on sensitive data}},
  author = {Wiktoria Leks},
  year = {2025},
  month = nov,
  note = {Submitted to Defensive Acceleration Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/wikigen-biological-safeguards-for-collaborative-aiml-on-sensitive-data-jpew}},
  url = {https://apartresearch.com/sprints/projects/wikigen-biological-safeguards-for-collaborative-aiml-on-sensitive-data-jpew}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026