Bacteria control gene expression through sigma (s) factors, proteins that recognise specific DNA sequences and initiate transcription. Although this mechanism is fundamental to life, our understanding of how s-factors identify and activate promoters—especially across different bacterial species—has been limited. The SiD (sigma ID) project set out to address this knowledge gap by combining synthetic biology, large-scale functional experiments, and artificial intelligence (AI).
The project generated a unique dataset comprising more than six million synthetic regulatory DNA sequences. These were tested using a high-resolution in vitro transcription platform with 17 s-factors from three different bacterial species. By determining which sequences produced RNA transcripts, the project identified both active promoters (positive functional data) and non-functional sequences (negative data). The negative dataset proved equally important, because knowing which sequences do not function as promoters dramatically improves the accuracy and robustness of deep-learning models.
This represents the largest functional s-factor dataset ever produced and provides an entirely new foundation for understanding the “regulatory grammar” of bacteria. Using these data, the project developed deep-learning models capable of predicting promoter strength, transcription start sites, and sigma-factor specificity. Importantly, the models generalise across species, addressing a long-standing challenge in bacterial genomics where promoter structures differ significantly between organisms.
These results enable the rational design of synthetic promoters with tailored expression strengths. This has broad applications in biotechnology, industrial fermentation, sustainable microbial production, and biomedical research. The project also demonstrated that AI models trained on systematic functional data—not only genome sequences—offer superior predictive and explanatory power. This highlights the long-term value of high-quality experimental negative and positive datasets for AI-driven biology.
As a lasting contribution, the project is developing a public web portal where the trained models will be accessible to researchers and industry. The portal will serve both as a promoter prediction tool and as a basis for genome-wide annotation of regulatory elements, improving future automated gene-regulation analysis and the design of engineered production strains.
The SiD project has provided new insights into bacterial gene regulation, generated large open datasets, trained researchers at the interface of synthetic biology and AI, and created a foundation for new national and international initiatives in minimal genomes, regulatory biology, and machine-learning-guided DNA design. The results will continue to benefit Norwegian and international biotechnology research well beyond the end of the project.
SiD-prosjektet har styrket Norges kompetanse innen syntetisk biologi og beregningsbasert biologi. Prosjektet har etablert nye eksperimentelle og beregningsmessige metoder for å avdekke hvordan bakterielle promotorer gjenkjennes av RNA-polymerase sigmafaktorer, og generert mer enn seks millioner unike DNA-sekvenser av promotorer og tre millioner RNA-transkripter. De resulterende datasett og dyplæringsmodeller utgjør en av de største ressursene innen forskning for prediksjon av promotorfunksjon.
Prosjektet har styrket det tverrfaglige samarbeidet mellom molekylærbiologi, biofysikk og kunstig intelligens, og etablert nye internasjonale partnerskap med forskningsmiljøer i Tsjekkia, Sverige, Spania og USA. Samarbeidene har fortsatt etter prosjektperioden og har allerede resultert i flere store oppfølgingsprosjekter, inkludert EU-prosjektene BlueTools, MetaExplore og Xtream, samt det norske KSP-prosjektet Threads.
SiD har bidratt til kompetansebygging gjennom opplæring av unge forskere i tverrfaglige metoder og gjennom utviklingen av åpne digitale infrastrukturer for promotorprediksjon og datadeling. Den nettbaserte plattformen som ble initiert i prosjektet, vil gi et tilgjengelig verktøy for både akademiske og industrielle brukere.
På lengre sikt forventes resultatene å påvirke bioteknologi, syntetisk biologi ved å muliggjøre prediktiv kontroll av genuttrykk og legge grunnlaget for rasjonell genomdesign. Dette vil styrke innovasjonskapasiteten innen syntetisk biologi og bidra til mer bærekraftige bioteknologiske løsninger, både nasjonalt og internasjonalt.
We propose an interdisciplinary data-driven experimental study, SiD, that will enable improved understanding of transcriptional regulation in bacteria. In SiD, we will develop a high-throughput microfluidics-based in vitro transcription method, coupled to in vivo screening efforts with DNA and RNA sequencing. The experimental efforts will allow us to determine the cis-acting DNA sequences for 21 sigma-factors, belonging to three bacterial species, Escherichia coli, Pseudomonas putida and Bacillus subtilis, at the single-nucleotide resolution. With the large-scale data, we will perform in silico data analysis, using statistics and machine learning, to develop algorithms that can infer sigma-factor-specific cis-binding motifs within bacterial promoters. The proposed methodology can be applied to any host of interest, and thus opens up the possibility to expand our knowledge on transcriptional regulation to a wide range of microorganisms. The generated big data will be stored on Norwegian nationwide Sigma2 platform, and the metadata will be made accessible to all interested parties based on the four principles of Findability, Accessibility, Interoperability and Reusability. A project-specific website will also be established following Open Science practices, making scientific research, data and dissemination output of SiD easily and openly accessible. To meet these ambitious goals, SiD requires an interdisciplinary approach, and therefore it integrates the fields of synthetic and computational biology (NTNU,Norway) with research partners from the fields of biophysics (NTNU, Norway), genomics (UNIBI, Germany), biochemistry (UU, Sweden) and synthetic biology (CSIC, Spain).