Automated Author Profile

Tuvi Etzion

Current S-Index

1.3

Sum of Dataset Indices for all datasets

Average Dataset Index per Dataset

0.7

Average Dataset Index per dataset

Total Datasets

2

Total datasets for this author

Average FAIR Score

73.1%

Average FAIR Score per dataset

Total Citations

1

Total citations to the author's datasets

Total Mentions

0

Total mentions of the author's datasets

S-Index Interpretation

S-Index Over Time

Cumulative Citations Over Time

Cumulative Mentions Over Time

Datasets

Prepared binned DNA data storage datasets for reconstruction benchmarking.

This repository includes datasets from the following publications. 1 Grass, R. N., Heckel, R., Puddu, M., Paunescu, D., & Stark, W. J. Robust chemical preservation of digital information on DNA in silica with error-correcting codes. Angewandte Chemie International Edition, 54, 8, 2552–2555 (2015)2 Erlich, Y. & Zielinski, D. DNA fountain enables a robust and efficient storage architecture. Science, 355, 6328, 950–954 (2017).3 Srinivasavaradhan, S. R., Gopi, S., Pfister, H. D. & Yekhanin S. Trellis BMA: Coded Trace Reconstruction on IDS Channels for DNA Storage. in 2021 IEEE International Symposium on Information Theory (ISIT), Melbourne, Australia, 2453–2458 (2021). The datasets are given in a binned format to enhance the reproducibility of the results presented in the paper. Bar-Lev, D., Orr, I., Sabary, O., Etzion T., & Yakkobi, E.   Scalable and robust DNA-based storage via coding theory and deep learning. 2024. Detailed description of the formatThe binned format was created using the binning step described in the paper ("Scalable and robust DNA-based storage via coding theory and deep learning"). Each cluster of reads appears in the file with a header followed by the reads. More specifically:The header consists of 2 lines, the first corresponds to the encoded sequence of the clusters, and the second is a line of 18x“*” that should be ignoredThe reads in the clusters are provided after the header, where each read is given in a separate lineEach cluster ends with two empty linesData processingTo ease the processing of our datasets, we also provide the following Python scripts (see https://github.com/itaiorr/Deep-DNA-based-storage)reads_preprocessor.py includes our preprocessing procedure for the raw reads. The procedure detects  and truncates the primers binning.py - parses the file of the binned reads and creates two Python dictionaries. In the first dictionary, each key is an encoded sequence,  and the value is a  list of the reads in the cluster. In the second dictionary the keys are the index and the value is a list of the reads in the cluster.

Authors

  • Daniella Bar Lev ;
  • Itai Orr ;
  • Omer Sabary ;
  • Tuvi Etzion ;
  • Eitan Yaakobi
0 Citations0 Mentions73% FAIR0.5 Dataset Index
10.5281/zenodo.142965872024

Prepared binned DNA data storage datasets for reconstruction benchmarking.

This repository includes datasets from the following publications. 1 Grass, R. N., Heckel, R., Puddu, M., Paunescu, D., & Stark, W. J. Robust chemical preservation of digital information on DNA in silica with error-correcting codes. Angewandte Chemie International Edition, 54, 8, 2552–2555 (2015)2 Erlich, Y. & Zielinski, D. DNA fountain enables a robust and efficient storage architecture. Science, 355, 6328, 950–954 (2017).3 Srinivasavaradhan, S. R., Gopi, S., Pfister, H. D. & Yekhanin S. Trellis BMA: Coded Trace Reconstruction on IDS Channels for DNA Storage. in 2021 IEEE International Symposium on Information Theory (ISIT), Melbourne, Australia, 2453–2458 (2021). The datasets are given in a binned format to enhance the reproducibility of the results presented in the paper. Bar-Lev, D., Orr, I., Sabary, O., Etzion T., & Yakkobi, E.   Scalable and robust DNA-based storage via coding theory and deep learning. 2024. Detailed description of the formatThe binned format was created using the binning step described in the paper ("Scalable and robust DNA-based storage via coding theory and deep learning"). Each cluster of reads appears in the file with a header followed by the reads. More specifically:The header consists of 2 lines, the first corresponds to the encoded sequence of the clusters, and the second is a line of 18x“*” that should be ignoredThe reads in the clusters are provided after the header, where each read is given in a separate lineEach cluster ends with two empty linesData processingTo ease the processing of our datasets, we also provide the following Python scripts (see https://github.com/itaiorr/Deep-DNA-based-storage)reads_preprocessor.py includes our preprocessing procedure for the raw reads. The procedure detects  and truncates the primers binning.py - parses the file of the binned reads and creates two Python dictionaries. In the first dictionary, each key is an encoded sequence,  and the value is a  list of the reads in the cluster. In the second dictionary the keys are the index and the value is a list of the reads in the cluster.

Authors

  • Daniella Bar Lev ;
  • Itai Orr ;
  • Omer Sabary ;
  • Tuvi Etzion ;
  • Eitan Yaakobi
1 Citation0 Mentions73% FAIR0.9 Dataset Index
10.5281/zenodo.142965882024