FASTQ Quality Control

Quality control with fastp

FASTQ Quality Control

Quality control with fastp

Loading WebAssembly runtime...

About the FASTQ Quality Control page

This page runs fastp 0.20.1 compiled to WebAssembly, giving you the full quality control and trimming workflow for single-end and paired-end FASTQ in the browser, with an HTML and JSON report and no server involved.

fastp is the first tool most people run on raw reads. It assesses read quality, detects and removes sequencing adapters, optionally trims low-quality tails, corrects overlapping bases in paired-end reads, and produces a report you can attach to a methods section. Defaults are sensible, so most runs need no arguments beyond the input and output files.

Because everything executes locally, human sequencing data never leaves your machine. This matters most for clinical and embargoed samples, where uploading raw reads to a third-party site is not acceptable.

What this tool does

Per-base quality assessment

Quality scores are plotted across the read length so you can see the characteristic decline toward the 3' end and decide whether trimming is needed.

Automatic adapter detection

Adapters are identified from the data without you specifying a sequence, which removes the most common source of poor alignment rates.

Quality and length filtering

Discard reads below a mean quality threshold, those with too many ambiguous bases, and those outside a minimum or maximum length, all with sensible defaults.

Paired-end handling

Full support for paired input, including merging overlapping pairs, base correction at the overlap, and writing unpaired or failed reads to separate files.

Poly-G and poly-X trimming

Strip the poly-G tails typical of two-colour Illumina chemistry, and poly-X tails common in PacBio and Nanopore data.

HTML and JSON reports

Generate a shareable interactive HTML report plus a machine-readable JSON summary, so QC becomes a documented step rather than a memory.

Example commands

fastp

Run default QC with automatic input detection. The sensible starting point for most data.

fastp -i input1.fq -I input2.fq -o out1.fq -O out2.fq

Process paired-end reads, keeping the two files synchronised.

fastp -i input1.fq -I input2.fq --merge --merged_out merged.fq

Merge reads whose pairs overlap into a single fragment, useful for assembly and amplicon work.

fastp -i input1.fq -I input2.fq --correction

Correct base calls in the overlap region of paired reads, improving consensus accuracy downstream.

fastp -a AGATCGGAAGAGC

Supply an explicit adapter sequence when the platform's chemistry is known and detection is not confident.

fastp -x --poly_x_min_len 10

Enable poly-X tail trimming, appropriate for PacBio and Nanopore reads.

Frequently asked questions

Are my FASTQ files uploaded to a server?

No. fastp runs as WebAssembly inside your browser tab. Raw reads, including human sequencing data, are processed on your own device and are never transmitted or stored.

What do the default quality thresholds do?

By default fastp requires a mean Phred quality of 15 or above, allows up to 40 percent of bases below quality 15, rejects reads with more than 5 ambiguous bases, and trims low-quality tails from the 3' end. These are conservative defaults that suit most short-read data.

Should I trim adapters manually?

Usually not. fastp detects adapters from the data by default, which is more reliable than hard-coding a sequence. Only specify one with -a if you know the library chemistry and detection looks wrong in the report.

How do I tell Phred 33 from Phred 64 encoding?

Most modern data is Phred 33, which fastp assumes by default. If your report shows implausible quality values or garbled quality characters, try the -6 flag to treat input as Phred 64.

Should I always merge paired-end reads?

Not always. Merging is valuable for amplicon and assembly work where you want a single consensus fragment, but for standard alignment you should keep reads paired so you can retain insert size and use proper-pair information.

Why do I have so many unpaired reads?

One mate failing quality or length filters causes the pair to be broken. Unpaired output files exist for this reason. A very high unpaired fraction usually indicates over-trimming, aggressive length filtering, or poor input quality worth revisiting.

How large a dataset can I process?

The file is loaded into browser memory, so practical limits are set by your device. Multi-gigabyte runs are better done with a local fastp install. For a quick QC check, use --reads_to_process to inspect a subset first.

Related tools

All processing happens locally in your browser. Nothing you upload is transmitted or stored, as described in the privacy policy.