INBAL PREUSS

 

 

 

 

 

 

 

 

 


My research focuses on novel approaches for storing digital information in synthetic DNA, developing theoretical, algorithmic, and system-level foundations for reliable encoding, storage, and retrieval. DNA-based data storage offers a promising solution for long-term, ultra-dense archival data storage, addressing the growing limitations of conventional media. I am particularly interested in combinatorial encoding schemes, error-resilient codes, and sequencing coverage models that enable scalable and cost-efficient DNA storage systems.

 


My research projects include:

  • Combinatorial shortmer encoding for DNA storage:
    This work introduces a novel encoding paradigm that expands the traditional four-base DNA alphabet into a combinatorial space built from carefully designed short DNA fragments (shortmers). By leveraging this approach, I demonstrate up to a 6.5-fold increase in logical data density compared with standard DNA storage methods, while maintaining near-zero reconstruction error in simulations and pilot implementations.
  • End-to-end DNA storage workflow design and decoding algorithms:
    I study and design complete DNA storage workflows, including encoding strategies, error-correction codes, and efficient reconstruction algorithms, under realistic synthesis and sequencing error types. These frameworks aim to bridge the gap between theoretical capacity and practical performance in real-world DNA storage systems.
  • Sequencing coverage and decoding reliability models:
    To support reliable data recovery from combinatorial DNA encodings, I develop theoretical models that quantify sequencing coverage requirements and provide practical guidelines for achieving error-free message reconstruction in experimental settings.
  • Error-correcting codes tailored to combinatorial DNA channels:
    My work also includes the development and analysis of coding schemes that address error patterns unique to combinatorial DNA storage. These constructions correct asymmetric errors that arise when specific shortmer components are missing from sequencing reads, improving decoding robustness and overall system reliability.
    My work has been published in high-impact journals, including a paper in Scientific Reports, and featured in special issues of IEEE Transactions on Molecular, Biological, and Multi-Scale Communications highlighting a decade of progress in DNA-based data storage. My research has also been presented at leading international conferences and seminars, including the IEEE International Symposium on Information Theory (ISIT) and Dagstuhl Seminars.

 


Through my research, I strive to advance the theory and practice of DNA-based archival storage by introducing new encoding paradigms, analytical tools, and system designs that push the boundaries of data density, reliability, and cost-effectiveness. I am always interested in interdisciplinary collaborations connecting computer science, information theory, and molecular systems.