De novo genome assembly — eukaryotes

A plant, animal or fungal genome assembled from scratch, with gene annotation grounded in data rather than ab initio prediction alone.

When a eukaryotic genome assembly is worth it

data2biology assembles plant, animal and fungal genomes de novo: from long reads (PacBio HiFi or Oxford Nanopore), ordered to chromosome level using Hi-C data, with gene annotation supported by RNA-Seq. Eukaryotic genomes are orders of magnitude larger than bacterial ones, contain introns and extensive repetitive regions, and often occur in several haploid copies. Assembling them is a project measured in weeks rather than hours and needs different tools and a different quality-control discipline. Building your own reference pays off when you plan sustained work on the species: trait mapping, population studies, expression analyses or evolutionary comparisons.

We have worked with compact genomes as well as polyploid and highly heterozygous species, where separating haplotypes is a challenge in itself. If you are unsure whether your data suffices for a meaningful assembly, we will assess that before the project starts.

How the data flows

Below we describe how we typically run this type of analysis. It is not a fixed procedure: the choice of tools depends on the organism, library type, sequencing depth and research question, and standards in bioinformatics change quickly. We agree the methodology with you before the project starts and update it as tools and reference databases evolve.

Input
Long reads (PacBio HiFi or ONT) · optionally Hi-C for scaffolding · optionally RNA-Seq and proteins from related species for annotation
Quality control and k-mer profiling
Read assessment plus estimation of genome size, heterozygosity and ploidy before assembly — this drives the decisions that follow.
NanoPlotJellyfishGenomeScope2Smudgeplot
Contig assembly
The assembler is matched to read type and heterozygosity. With HiFi data, haplotype separation is possible at this stage.
hifiasmFlyeVerkkoCanu
Removing haplotypic duplication
Redundant copies of heterozygous regions inflate assembly size and artificially raise the count of duplicated genes. We remove them based on coverage distribution.
purge_dups
Scaffolding to chromosome level
Hi-C data lets us order and orient contigs into pseudochromosomes. We verify the result on the contact map and curate it manually where needed.
YaHSSALSA2PretextViewJuicebox
Assembly quality assessment
Contiguity, completeness of single-copy gene sets, and an independent accuracy estimate from read k-mers.
QUASTBUSCOMerqury
Repeat identification and masking
We build a species-specific repeat library rather than relying only on generic databases, then soft-mask the genome before gene annotation.
RepeatModeler2RepeatMaskerEDTA
Structural gene annotation
Gene models are built from evidence: RNA-Seq transcripts and proteins from related species, supported by ab initio prediction. tRNAs and other ncRNAs are annotated separately.
BRAKER3HelixerfunannotatetRNAscan-SEInfernal
Functional annotation
Assignment of domains, GO terms, KEGG pathways and product descriptions from similarity to protein databases and domain signatures.
InterProScaneggNOG-mapperDIAMOND vs Swiss-Prot
Data package and genome browser
We configure a browser where you can inspect any gene alongside the evidence tracks, and prepare submission-ready files.
JBrowse 2IGV
Output
Genome sequence · GFF3 annotation · proteins and CDS · repeat library · BUSCO and Merqury reports · functional annotation spreadsheet
A typical data flow for eukaryotic genome assembly. Step 4 is performed only when Hi-C data is available.

What we pay attention to

Material quality decides the outcome

Long-read assembly requires high molecular weight DNA. Degraded material yields short reads, and short reads yield a fragmented assembly that no amount of computation can repair. If extraction is still ahead of you, that conversation is better had sooner than later.

How to read a BUSCO result

High completeness does not automatically mean a good assembly, and a high proportion of duplicated genes is often a sign of unresolved haplotypic redundancy rather than genuine genome duplication. We therefore report BUSCO alongside Merqury and contiguity statistics rather than as a single number out of context.

Annotation without RNA-Seq data

Annotation can be done from proteins of related species alone, but the gene models are then noticeably less accurate — particularly at exon boundaries, alternative isoforms and untranslated regions. Where possible, sequencing even a handful of RNA libraries from different tissues or conditions is one of the highest-return investments in the whole project.

What you receive

File Format What it contains
genome.fasta FASTA The assembled genome sequence. With Hi-C scaffolding the records correspond to pseudochromosomes, with remaining sequences in an unplaced pool.
genome.softmasked.fasta FASTA The same sequence with repetitive regions in lower case — the version used for annotation and mapping.
annotation.gff3 GFF3 Gene models in a hierarchy: gene contains mRNA, and mRNA contains exons, CDS and UTR regions. One line per feature, nine columns.
  • seqid, source, type — chromosome, model source, feature kind
  • start, end, strand, phase — coordinates, orientation and reading frame
  • attributes — ID and Parent linking features into the hierarchy, plus name and description
  • A measure of transcript-evidence support for each model, where RNA-Seq data was available
proteins.faa / cds.fna FASTA Protein and coding sequences for all gene models, with identifiers matching the GFF3 file.
functional_annotation.xlsx XLSX The gene catalogue as a spreadsheet — the most frequently opened file in the whole package.
  • gene_id, transcript_id, chromosome, coordinates, exon count, length
  • Best Swiss-Prot match with e-value and identity
  • Pfam domains and InterPro signatures
  • GO terms, KEGG KO numbers and COG categories
  • Product description and a flag for RNA-Seq support of the model
repeats/ FASTA + GFF3 + TSV A species-specific repeat library, repeat locations in the genome, and a summary of the share of each element class.
quality/ HTML + TXT Assembly quality reports, presented together rather than in isolation.
  • QUAST: scaffold count, N50, L50, total length, gap fraction
  • BUSCO: complete single-copy, duplicated, fragmented and missing
  • Merqury: QV value and k-mer completeness against the reads
  • Hi-C contact map, where scaffolding was performed
jbrowse2/ configuration A ready genome browser configuration with annotation, repeat and RNA-Seq coverage tracks — to run locally or on your own server.
methods.docx DOCX A methods section draft with tool versions, parameters, assembly statistics and citations.

What we need from you

  • Long reads in FASTQ or BAM format, with platform and chemistry details
  • Estimated genome size and ploidy, if known from cytometry or the literature
  • Hi-C data, if you expect a chromosome-level assembly
  • RNA-Seq data from as many tissues or conditions as possible — it markedly improves annotation
  • Related species with available genomes that can be used as supporting evidence

Quotation and how we work

We quote every project individually. The cost depends mainly on the number of samples, the availability of a reference genome, the number of comparisons and the scope of analyses you ask for.

  1. You tell us what you want to find out and what data you have.
  2. If anything needs clarifying we arrange a short call — free of charge and without obligation.
  3. Within 3–5 working days you receive a quote with the scope of work, the list of result files and a delivery date.
  4. You decide. The quote carries no obligation.

The scope and cost of projects like these vary widely enough that quoting a single representative figure would be misleading. If you need a ballpark before a formal quote, just ask — we can give a rough range straight away.

Any figures given are indicative and net of tax; VAT is added according to the applicable regulations. They do not constitute a binding offer. More about how we work in the FAQ.

Got data to analyse?

Tell us what you want to find out and we will scope the analysis and timeline together. Discussing the project and preparing a quote are free of charge.

Request a quote See a worked example