RNA-Seq analysis
From raw FASTQ files to a gene list that actually answers your research question.
When RNA-Seq is the right tool
data2biology carries out complete RNA-Seq analyses: from raw FASTQ files to a list of differentially expressed genes with GO and KEGG enrichment. A project covering 10 to 20 samples usually takes about four weeks and costs approximately PLN 4,500 net (as of 2026). RNA-Seq answers the question of which genes change activity between conditions: treatment versus control, tissue versus tissue, timepoint versus timepoint, mutant versus wild type. Beyond the gene list itself, the analysis provides functional context — which biological processes and pathways are over-represented among those changes.
We work both with model organisms that have a well-annotated reference genome and with non-model species, where the starting point is often a de novo assembled transcriptome. We handle single- and paired-end data, stranded and unstranded libraries, and designs with several factors or time series.
How the data flows
Below we describe how we typically run this type of analysis. It is not a fixed procedure: the choice of tools depends on the organism, library type, sequencing depth and research question, and standards in bioinformatics change quickly. We agree the methodology with you before the project starts and update it as tools and reference databases evolve.
What we pay attention to
Design and replication
The number of biological replicates affects the outcome more than the choice of software. With three replicates per group we detect clear changes, but weaker effects stay out of statistical reach. If the experiment is still being planned, we are glad to discuss group sizes and sequencing depth before any data exists — that is the cheapest moment for that conversation.
Significance thresholds
By default we treat genes with an adjusted p-value below 0.05 and an absolute log2 fold change of at least 1 as differentially expressed. These thresholds are a convention rather than a law of nature — if your field uses different ones we apply those, and we always report which values were used.
Non-model organisms
Without a reference genome we start from a de novo transcriptome assembly and functional annotation by similarity to protein databases. Interpretation is then more cautious, because some transcripts remain without a reliable functional assignment; we note this in the report.
What you receive
We deliver results in formats you can open without any bioinformatics background, and in parallel in formats suitable for further programmatic analysis. Below is an example file set for a project with a single comparison — the actual list depends on the scope and is agreed in the quote.
To see what such a package looks like in practice, have a look at our worked example: an RNA-Seq analysis of the Arabidopsis drought response — built on simulated data, but in exactly the format and scope a real project delivers.
| File | Format | What it contains |
|---|---|---|
| multiqc_report.html | HTML |
Aggregated quality report for all samples — opens in a browser.
|
| counts_raw.tsv | TSV |
Raw count matrix: rows are genes, columns are samples. The starting point for any re-analysis.
|
| expression_TPM.xlsx | XLSX |
Normalised expression levels, ready to browse in Excel.
|
| DE_results_<contrast>.xlsx | XLSX |
The main result file: the full differential expression table, one sheet per comparison.
|
| enrichment_GO_KEGG.xlsx | XLSX |
Enrichment results, separately for up- and down-regulated genes.
|
| figures/ | PNG + PDF |
Figures at publication resolution (PNG) and as vectors (PDF/SVG) for further editing.
|
| alignments/*.bam + *.bai | BAM | Aligned reads with indexes — you can inspect any gene in a genome browser such as IGV. |
| tracks/*.bw | bigWig | Library-size normalised coverage tracks for IGV or JBrowse. |
| methods.docx | DOCX | A ready draft of the Materials and Methods section with all tool versions, parameters and citations — for use in your manuscript. |
What we need from you
- Raw FASTQ files (gzip-compressed), ideally straight from the sequencer with no prior trimming
- A sample table: file name, experimental group, replicate, and any batch information
- The species and reference genome version if you have a preference — otherwise we select a current one
- A description of your research question and the comparisons you care about
- Library details: single-end or paired-end, stranded or not, poly(A) or rRNA-depleted
We agree the transfer method individually — most often we receive data on physical media or set up a dedicated SFTP account. Data is processed on our own servers and on servers rented from established providers.
Quotation and how we work
We quote every project individually. The cost depends mainly on the number of samples, the availability of a reference genome, the number of comparisons and the scope of analyses you ask for.
- You tell us what you want to find out and what data you have.
- If anything needs clarifying we arrange a short call — free of charge and without obligation.
- Within 3–5 working days you receive a quote with the scope of work, the list of result files and a delivery date.
- You decide. The quote carries no obligation.
For orientation: a standard RNA-Seq analysis of 10–20 samples, from quality control to a differential gene list with GO and KEGG enrichment, typically costs around PLN 4,500 net (approx. EUR 1,100). That assumes an organism with an available reference genome and annotation, one or two main comparisons, and the result file set described above. De novo transcriptome assembly, isoform and splicing analyses, multifactorial designs and time series are quoted separately.
Any figures given are indicative and net of tax; VAT is added according to the applicable regulations. They do not constitute a binding offer. More about how we work in the FAQ.
Got data to analyse?
Tell us what you want to find out and we will scope the analysis and timeline together. Discussing the project and preparing a quote are free of charge.
Request a quote See a worked example