Talk to Our Scientists
"*" indicates required fields
De Novo sequencing at CovalX
De novo sequencing involves reconstructing the amino acid sequence of a protein directly from the MS/MS spectrum, without a reference. The term “de novo” literally means “from the beginning”: de novo sequencing starts from scratch, from the raw signal.
CovalX applies multi-enzyme proteolytic digestion combined with nano-flow liquid chromatography coupled to high-resolution MS/MS fragmentation to reconstruct the full amino acid sequence from fragment ion data. The approach is used across antibody sequencing programs, biosimilar characterization, and sequence verification workflows where database-dependent methods are insufficient.
De Novo sequencing workflow
- Quality Control. The protein is submitted to intact HM-MALDI MS analysis to verify the integrity and purity of the protein sample to sequence.
- Multi-enzyme proteolytic digestion. Multiple proteolytic enzymes are applied in parallel. Each enzyme generates a distinct set of peptides; the overlapping coverage across all digests is what allows complete sequence reconstruction without gaps.
- Nano-flow LC-MS/MS. Peptide pools are separated by nano-flow liquid chromatography and introduced into the mass spectrometer. Each peptide is fragmented by HCD, generating b- and y-ion series that encode the amino acid sequence. Where isobaric residues are present (leucine and isoleucine, asparagine and GlyGly), a second fragmentation step, EThcD (electron-transfer/higher-energy collisional dissociation), resolves the ambiguity at the fragment ion level.
- Sequence assembly. Dedicated de novo software aligns overlapping peptide sequences across all enzyme digests to reconstruct the full protein sequence.
Want to know more?
Our team will help you design your experiment
De Novo Sequencing Deliverables and Report Structure
| Intact QC data | High-Mass MALDI MS spectra confirming your molecule integrity and purity. Delivered as part of the project report. |
|---|---|
| Full amino acid sequence | Reconstructed from overlapping peptide coverage across all enzyme digests. |
| Peptide mass fingerprint | Per enzyme digest, with a sequence coverage map showing which regions are covered by which proteolytic set. The report includes an overlap map for each chain, plus a summary table of coverage percentage, chain length, and total peptides identified. |
| PTM profiling | Main post-translational modifications identified at the peptide level during sequencing, including deamidation, oxidation, and pyroglutamate formation. Each modification is reported with its exact position on the sequence. |
| Glycan profiling | N-glycan site position, occupancy, and structure, determined from glycopeptide analysis performed on the same sample. Glycan species are reported with their relative abundance. |
| Full analytical report | Written report covering materials and methods, coverage per enzyme, and interpretation. Structured for direct use in IND, BLA, or biosimilar dossiers. |
Why Choose CovalX for De Novo Sequencing?
Intact mass analysis as a systematic first step
Before digestion begins, CovalX performs High-Mass MALDI intact mass analysis on the protein to sequence. This step confirms molecular weight, detects unexpected variants or truncations, and identifies heterogeneity that would otherwise propagate through the sequencing workflow undetected.
Glycan profiling included
For glycoproteins and antibodies, CovalX analyzes glycopeptides by high-resolution MS/MS. This analysis determines N-glycan site position, site occupancy, and structure. The data is delivered alongside the sequencing results as part of the project report. For programs requiring both sequence confirmation and glycan characterization, both datasets are generated from a single sample submission without additional material.
PTM profiling included
The same high-resolution MS/MS analysis used for sequencing also identifies main post-translational modifications (PTMs) at the peptide level, including their site and relative abundance. This data is delivered alongside the sequencing results as part of the project report. No additional sample submission is required.
EThcD fragmentation for isobaric residue disambiguation
Leucine and isoleucine are isobaric by standard hcD fragmentation. CovalX applies EThcD (electron-transfer/higher-energy collisional dissociation) systematically where isobaric residues are detected, resolving ambiguities at the fragment ion level. This is relevant for programs where sequence accuracy is required for INN applications or biosimilar dossiers.
No database submission, full confidentiality
Sequence data never enters a public database. This is a hard requirement for pharmaceutical clients with proprietary sequences.
Direct access to scientists with over two decades of MS experience
Projects are handled by scientists with specific expertise in de novo MS sequencing. Sample suitability, enzyme selection strategy, and delivery timelines are evaluated before.
Reference-free protein sequencing
Sequence your protein of interest
Technical Notes
CovalX runs de novo sequencing on high-resolution nano-LC coupled to mass spectrometers capable of HCD and EThcD fragmentation. Key parameters:
- Multi-enzyme digestion protocol to maximize sequence coverage.
- Nano-flow LC separation: optimized for peptide resolution and sensitivity at the sample quantities used.
- HCD fragmentation for standard peptide sequencing; EThcD applied for isobaric residue differentiation.
- De novo assembly software aligns overlapping peptide sequences across enzyme sets to reconstruct the full protein sequence.
Sample requirements and material compatibility (including buffer composition and purity expectations) are discussed at project setup. Samples should be provided as purified protein; CovalX performs concentration or gel-based separation as needed prior to digestion.
Regulatory Context
De novo sequencing is used as the primary sequence determination method when gene sequencing is unavailable, incomplete, or restricted. In IND and BLA submissions, primary sequence confirmation is a required element of structural characterization under ICH Q6B. For biosimilar programs, sequence confirmation by an analytical method independent of the reference product record supports the comparability case required by FDA and EMA.
De novo MS sequencing data provides residue-level sequence evidence that complements gene sequencing where available, or substitutes for it where the genomic record is absent. The method produces data directly from the drug substance, which is the regulatory requirement.
Frequently Asked Questions
What is de novo protein sequencing by mass spectrometry?
De novo MS sequencing reconstructs the amino acid sequence of a protein directly from MS/MS fragment ion spectra, without consulting a reference database. The sequence is read from the mass differences between consecutive fragment ions in the b- and y-ion series generated during peptide fragmentation.
How long does a De Novo sequencing project take?
Standard projects are delivered within 1 to 2 weeks from sample receipt. Expedited timelines are available on request.
What sample quantity is required?
Standard projects require 100 µg of purified protein or antibody. Final requirements depend on material purity and the number of enzyme digests planned, and are confirmed at project setup.
What are the deliverables?
Clients receive the full reconstructed amino acid sequence including CDRs positions (in case of mAb sequencing), per-enzyme peptide coverage maps, PTMs and N-Glycan analysis in a structured analytical report for regulatory use.
Which regulatory submissions use de novo sequencing data?
Primary sequence confirmation is required under ICH Q6B for biotherapeutic characterization packages in IND and BLA submissions. De novo sequencing provides this confirmation for novel molecules lacking a verified genomic sequence record. In biosimilar dossiers, it supports the sequence identity claim required by FDA and EMA.

CovalX De Novo Sequencing Brochure