De novo genome assembly — eukaryotes
A plant, animal or fungal genome assembled from scratch, with gene annotation grounded in data rather than ab initio prediction alone.
When a eukaryotic genome assembly is worth it
data2biology assembles plant, animal and fungal genomes de novo: from long reads (PacBio HiFi or Oxford Nanopore), ordered to chromosome level using Hi-C data, with gene annotation supported by RNA-Seq. Eukaryotic genomes are orders of magnitude larger than bacterial ones, contain introns and extensive repetitive regions, and often occur in several haploid copies. Assembling them is a project measured in weeks rather than hours and needs different tools and a different quality-control discipline. Building your own reference pays off when you plan sustained work on the species: trait mapping, population studies, expression analyses or evolutionary comparisons.
We have worked with compact genomes as well as polyploid and highly heterozygous species, where separating haplotypes is a challenge in itself. If you are unsure whether your data suffices for a meaningful assembly, we will assess that before the project starts.
How the data flows
Below we describe how we typically run this type of analysis. It is not a fixed procedure: the choice of tools depends on the organism, library type, sequencing depth and research question, and standards in bioinformatics change quickly. We agree the methodology with you before the project starts and update it as tools and reference databases evolve.
What we pay attention to
Material quality decides the outcome
Long-read assembly requires high molecular weight DNA. Degraded material yields short reads, and short reads yield a fragmented assembly that no amount of computation can repair. If extraction is still ahead of you, that conversation is better had sooner than later.
How to read a BUSCO result
High completeness does not automatically mean a good assembly, and a high proportion of duplicated genes is often a sign of unresolved haplotypic redundancy rather than genuine genome duplication. We therefore report BUSCO alongside Merqury and contiguity statistics rather than as a single number out of context.
Annotation without RNA-Seq data
Annotation can be done from proteins of related species alone, but the gene models are then noticeably less accurate — particularly at exon boundaries, alternative isoforms and untranslated regions. Where possible, sequencing even a handful of RNA libraries from different tissues or conditions is one of the highest-return investments in the whole project.
What you receive
| File | Format | What it contains |
|---|---|---|
| genome.fasta | FASTA | The assembled genome sequence. With Hi-C scaffolding the records correspond to pseudochromosomes, with remaining sequences in an unplaced pool. |
| genome.softmasked.fasta | FASTA | The same sequence with repetitive regions in lower case — the version used for annotation and mapping. |
| annotation.gff3 | GFF3 |
Gene models in a hierarchy: gene contains mRNA, and mRNA contains exons, CDS and UTR regions. One line per feature, nine columns.
|
| proteins.faa / cds.fna | FASTA | Protein and coding sequences for all gene models, with identifiers matching the GFF3 file. |
| functional_annotation.xlsx | XLSX |
The gene catalogue as a spreadsheet — the most frequently opened file in the whole package.
|
| repeats/ | FASTA + GFF3 + TSV | A species-specific repeat library, repeat locations in the genome, and a summary of the share of each element class. |
| quality/ | HTML + TXT |
Assembly quality reports, presented together rather than in isolation.
|
| jbrowse2/ | configuration | A ready genome browser configuration with annotation, repeat and RNA-Seq coverage tracks — to run locally or on your own server. |
| methods.docx | DOCX | A methods section draft with tool versions, parameters, assembly statistics and citations. |
What we need from you
- Long reads in FASTQ or BAM format, with platform and chemistry details
- Estimated genome size and ploidy, if known from cytometry or the literature
- Hi-C data, if you expect a chromosome-level assembly
- RNA-Seq data from as many tissues or conditions as possible — it markedly improves annotation
- Related species with available genomes that can be used as supporting evidence
Quotation and how we work
We quote every project individually. The cost depends mainly on the number of samples, the availability of a reference genome, the number of comparisons and the scope of analyses you ask for.
- You tell us what you want to find out and what data you have.
- If anything needs clarifying we arrange a short call — free of charge and without obligation.
- Within 3–5 working days you receive a quote with the scope of work, the list of result files and a delivery date.
- You decide. The quote carries no obligation.
The scope and cost of projects like these vary widely enough that quoting a single representative figure would be misleading. If you need a ballpark before a formal quote, just ask — we can give a rough range straight away.
Any figures given are indicative and net of tax; VAT is added according to the applicable regulations. They do not constitute a binding offer. More about how we work in the FAQ.
Got data to analyse?
Tell us what you want to find out and we will scope the analysis and timeline together. Discussing the project and preparing a quote are free of charge.
Request a quote See a worked example