OpenCRISPR-1 is the first gene editor designed entirely by artificial intelligence and validated in human cells, a fully synthetic Cas9-like protein that does not exist in nature. Published in Nature in 2025, it is open source, free to license for ethical research and commercial use, and represents a meaningful alternative to SpCas9 for applications where off-target fidelity and IP access are limiting factors.
Why SpCas9 still carries limitations after years of use
SpCas9 was discovered in Streptococcus pyogenes and adapted for human genome editing through foundational work in the early 2010s. It remains the most widely used CRISPR nuclease, but it is a borrowed tool, shaped by bacterial evolutionary pressures rather than human therapeutic requirements. Several limitations persist in clinical and manufacturing contexts:
- 1 Off-target editing SpCas9 exhibits measurable cleavage at unintended genomic loci, raising genotoxicity concerns. In clinical manufacturing, this requires extensive off-target screening and adds risk to IND applications.
- 2 Pre-existing immunity Studies have detected anti-SpCas9 antibodies and T-cell responses in a substantial proportion of healthy donors, owing to its origin in a common human pathogen. This limits in vivo delivery strategies and complicates patient selection in clinical trials.
- 3 AAV payload constraint At ~4.2 kb, SpCas9 occupies most of the practical packaging capacity of AAV vectors (~4.5–4.7 kb usable). This leaves minimal room for regulatory elements, limiting in vivo delivery flexibility.
- 4 IP access and licensing The SpCas9 patent landscape involves multiple overlapping claims across academic institutions and commercial entities. Obtaining freedom-to-operate for therapeutic programs is time-consuming and expensive.
How Profluent built a gene editor from a language model
Profluent Bio, a San Francisco-based AI research company, used a protein language model to generate OpenCRISPR-1 from scratch. Rather than engineering variants of existing nucleases, they trained a model on natural CRISPR diversity and sampled entirely new sequences from the learned distribution. The development process moved from curated genomic data through model training, generative sampling, and functional screening in human cells:
CRISPR-Cas Atlas (~26 TB)
A curated database of assembled genomes and metagenomes. 238,917 Cas9-like protein sequences were selected as high-quality training data for the model.
Pre-training ProGen2 on Cas sequences
The protein language model learned sequence co-evolution patterns, structural constraints, and functional signatures embedded in natural CRISPR-Cas diversity — without being given explicit structural or mechanistic rules.
Generative sampling of novel sequences
The trained model generated millions of Cas9-like protein sequences not found in nature. These are not mutations of SpCas9, but entirely new sequences sampled from the learned distribution of functional Cas9-like proteins.
Functional screening in HEK293T cells
209 candidates were characterized for editing activity via plasmid delivery. 48 were selected for rigorous multi-site on- and off-target analysis. OpenCRISPR-1 was the top-performing sequence.
AI-generated guide RNA
A separate model generated a compatible synthetic sgRNA for OpenCRISPR-1. Testing showed the AI-generated gRNA improved editing efficiency over standard Cas9 gRNAs in 4 of 5 nucleases tested.
Performance data and molecular properties
OpenCRISPR-1 retains the prototypical Type II Cas9 architecture: RuvC domain, HNH domain, NGG PAM preference. However, its amino acid sequence is entirely novel. Current published data comes from plasmid-based delivery in HEK293T cells; RNP-format and primary cell characterization are ongoing.
| Property | OpenCRISPR-1 | SpCas9 | Status |
|---|---|---|---|
| PAM requirement | NGG | NGG | Confirmed |
| On-target efficiency (HEK293T) | 55.7% | 48.3% | Confirmed |
| Off-target editing rate | 0.32% | 6.1% | Confirmed |
| Base editing compatibility (A-to-G) | Yes (ABE8.20, PF-DEAM-1/2) | Yes | Confirmed |
| Compatible with canonical SpCas9 gRNAs | Yes | Yes | Confirmed |
| RNP delivery characterization | Ongoing | Extensive | In progress |
| Genome-wide specificity (unbiased) | Ongoing | Extensive | In progress |
| In vivo immunogenicity data | Not yet published | Substantial | Pending |
| Licensing | Open source, free (ethical use) | Complex, multi-party | Open license |
Application in base editing: OpenCRISPR-1 has been validated in a base editing architecture using both ABE8.20 (an established engineered deaminase) and two AI-generated deaminases — PF-DEAM-1 and PF-DEAM-2 — developed by a separate Profluent-trained model. Activity and specificity were maintained or improved relative to SpCas9 base editing at matched target sites.
Where OpenCRISPR-1 is most relevant today
CAR-T and TCR-T manufacturing
Ex vivo T-cell engineering is one of the most immediate use cases. Higher on-target efficiency and reduced off-target activity could improve editing yield and the safety profile of gene-edited T-cell products. Open licensing removes a common commercial barrier for CDMOs and cell therapy developers.
iPSC engineering
Unintended edits introduced during pluripotent cell engineering carry forward through differentiation and can confound downstream phenotyping. The reduced off-target rate is particularly valuable here, where clonal expansion amplifies the consequences of any mis-edit.
Single-nucleotide correction
Validated A-to-G base editing compatibility makes OpenCRISPR-1 relevant for programs targeting point mutations, such as those in hemoglobin disorders, metabolic conditions, and other monogenic diseases addressable by base correction rather than full cleavage.
Modular base editing platform development
The availability of AI-generated deaminase fusions (PF-DEAM-1, PF-DEAM-2) alongside OpenCRISPR-1 opens the door to constructing base editors where every component is AI-designed, potentially enabling further optimization through iterative AI-driven protein engineering.
Programs with IP constraints
For academic labs or early-stage biotech teams where SpCas9 licensing is a practical obstacle, OpenCRISPR-1's free open-source license (requiring only a signed ethical use agreement for commercial applications) substantially lowers the barrier to entry.
What you should know before adopting OpenCRISPR-1
OpenCRISPR-1 is a promising and well-characterized research tool, but several data gaps remain important for researchers planning therapeutic applications or head-to-head comparisons:
Genome-wide specificity data
While SITE-Seq was used for OpenCRISPR-1 validation, other unbiased approaches such as GUIDE-seq, Digenome-seq, or CIRCLE-seq across a broad genomic panel are generally required for comprehensive off-target identification. Plasmid-based HEK293T data is encouraging but not a substitute.
RNP format performance
All published efficiency data uses plasmid delivery. RNP characterization (the preferred format for clinical ex vivo editing) is described as ongoing by Profluent.
Primary cell and in vivo data
Editing efficiency and specificity in primary HSCs, T cells, and hepatocytes, as well as in vivo pharmacokinetics, have not yet been published. Behavior in differentiated primary cells may differ substantially from HEK293T.
Immunogenicity profile
The hypothesis that a non-pathogen-derived sequence may reduce pre-existing immunity is biologically plausible but unconfirmed. Rigorous clinical immunogenicity assessments still need to be performed and reported.
Licensing for commercial use
While research use is free, commercial therapeutic applications require signing Profluent's license agreement, which includes ethical use obligations and excludes human germline editing.
IP status
Profluent has filed IP on OpenCRISPR-1. While access is free under their open license, the sequence is not patent-free. Researchers developing therapeutic products should review the license terms carefully.
A new mode of tool development for biology
What OpenCRISPR-1 represents is not just a new tool, but a demonstration that AI can now operate as a primary author of functional biology. The CRISPR-Cas Atlas, assembled across 26 TB of genomic sequence, gave a language model enough biological context to generate proteins that edit the human genome as well as or better than a tool that took decades of natural evolution and laboratory optimization to produce.
Since the open-source release in April 2024, tens of thousands of researchers across therapeutics, agriculture, and drug discovery have accessed the sequence. The Nature 2025 publication formalizes the platform's scientific basis and expands the characterization data available for evaluation.
OpenCRISPR-1 is best understood right now as a high-performing, IP-accessible research nuclease with an unusually strong specificity profile in cell line data, and a growing body of evidence for base editing compatibility. The gaps in RNP data, primary cell performance, and in vivo immunogenicity are real, and researchers planning clinical programs will want to generate that data in their own hands. But for ex vivo research workflows, functional screens, and programs where SpCas9 licensing is a barrier, OpenCRISPR-1 is a practical and principled choice to evaluate.
- Ruffolo J, Nayfach S, Gallagher J, Bhatnagar A, et al. "Design of highly functional genome editors by modelling CRISPR-Cas sequences." Nature, 2025.
- Profluent Bio. OpenCRISPR initiative. profluent.bio/modality/opencrispr. Profluent press release. Business Wire, April 22, 2024.
- GEN — Genetic Engineering & Biotechnology News. "Profluent's AI-Designed Gene Editor Glimpses into Generalizable Platform." August 2025.