What it is
Engineering proteins that recognize an arbitrary chosen DNA sequence has been an unsolved design problem, with past progress limited to reprogramming natural binders through selection. This work describes a computational method that designs small DNA-binding proteins reading short target sequences through major-groove contacts, and uses it to generate binders for five distinct DNA targets with mid- to high-nanomolar affinities. Individual modules match the computational specificity model at as many as six base-pair positions, and can be rigidly arrayed along the double helix with RFdiffusion for higher-order specificity. A crystal structure of a designed complex closely matched the design model, and the proteins repressed and activated neighboring genes in both E. coli and mammalian cells.
Why it matters
Sequence-specific DNA readers underpin genome editing and gene regulation, and existing tools (zinc fingers, TALEs, CRISPR) each carry constraints on delivery and design. Building small binders from scratch that match the intended specificity at up to six base pairs offers a programmable alternative that is compact and therefore easier to deliver. Function in both bacterial and mammalian cells shows the designs work as regulators inside living systems, not only in vitro.
Underlined numbers link to their source. Every metric and quoted figure is listed under Sources and data below.
Filed underprotein design, DNA binding, gene regulation, synthetic biology, genome editing