Skip to content

Initial upload of CVA10 dataset - #466

Draft
nneune wants to merge 16 commits into
masterfrom
add-cva10
Draft

nneune wants to merge 16 commits into
masterfrom
add-cva10

Conversation

@nneune

@nneune nneune commented Jul 21, 2026

Copy link
Copy Markdown
Collaborator

Description of proposed changes

This PR introduces a new Coxsackievirus A10 dataset rooted on the Kowalik reference sequence. It includes all essential files for Nextclade compatibility, such as the reference genome, genome annotation, and detailed documentation. The dataset supports subgenotype assignment, phylogenetic placement, and sequence quality control for Coxsackievirus A10.

New Coxsackievirus A10 Dataset

  • Dataset documentation:

    • Added a comprehensive README.md explaining dataset scope, subgenogroup definitions, reference types, and usage instructions.
    • Added a CHANGELOG.md with an initial release entry.
  • Reference data and genome annotation:

    • Added the Kowalik reference genome sequence in reference.fasta.
    • Added a GFF3 genome annotation file with gene and protein features for the Kowalik reference.

Checklist

  • Send out to ENPEN members to test

@nneune
nneune had a problem deploying to refs/pull/466/merge July 21, 2026 09:45 — with GitHub Actions Failure
@nneune
nneune had a problem deploying to refs/heads/add-cva10 July 22, 2026 08:44 — with GitHub Actions Error
@nneune
nneune temporarily deployed to refs/pull/466/merge July 22, 2026 08:45 — with GitHub Actions Inactive
@nneune
nneune had a problem deploying to refs/heads/add-cva10 July 22, 2026 10:14 — with GitHub Actions Error
@nneune
nneune temporarily deployed to refs/pull/466/merge July 22, 2026 10:14 — with GitHub Actions Inactive
nextstrain-bot and others added 2 commits July 22, 2026 10:16
…on the inferred reference sequence

Previous builds inferred the tree's root sequence independently of the
dataset reference, causing mismatched mutations at the root and a
Nextclade preprocessing error. The tree is now rebuilt with ancestral
reconstruction anchored to the reference (inferred ancestral) sequence.
@nneune
nneune had a problem deploying to refs/heads/add-cva10 July 22, 2026 13:17 — with GitHub Actions Error
@nneune
nneune temporarily deployed to refs/pull/466/merge July 22, 2026 13:17 — with GitHub Actions Inactive
@nneune
nneune had a problem deploying to refs/heads/add-cva10 August 3, 2026 10:36 — with GitHub Actions Error
@nneune
nneune deployed to refs/pull/466/merge August 3, 2026 10:36 — with GitHub Actions Active
nextstrain-bot and others added 2 commits August 3, 2026 10:37
- change order of references in Nextclade results table
- specify correct order in README
- remove non-VP1 mutations for VP1-only clades (previous issue fixed)
@nneune
nneune deployed to refs/heads/add-cva10 September 15, 2026 16:27 — with GitHub Actions Active
@nneune
nneune deployed to refs/pull/466/merge September 15, 2026 16:34 — with GitHub Actions Active
@nneune

nneune commented Sep 15, 2026

Copy link
Copy Markdown
Collaborator Author

@ivan-aksamentov
ivan-aksamentov deployed to refs/pull/466/merge September 15, 2026 17:02 — with GitHub Actions Active
@ivan-aksamentov

Copy link
Copy Markdown
Member

Note

✨ AI is responding on behalf of Ivan

First review of PR #466, adding the enpen/enterovirus/cva10 dataset, at head fefe1a7d against base ebb92c18.

Verdict: no blocking issues. The dataset is structurally sound and scientifically well founded: coordinates, CDS divisibility, GFF3/tree agreement, generated output, registration, and a live Nextclade CLI run all pass. There are 9 non-blocking items (5 medium, 4 low), all concerning documentation accuracy, user-facing links, and reference-tree presentation rather than analysis correctness.

Testing

Observed

What the dataset is, and where it diverges from its description [click to expand]

The PR adds one whole-genome Coxsackievirus A10 dataset at enpen/enterovirus/cva10, registers it in data/enpen/collection.json between cva16 and ev-a71, and ships the generated output under data_output/. It follows the pattern of the released cva16 and ev-a71 siblings: reference.fasta holds a reconstructed ancestral sequence (>ancestral_sequence, 7409 nt) laid out on the coordinate system of the Kowalik prototype (AY421767.1), the GFF3 carries Kowalik's mature-peptide coordinates, and Kowalik is kept as a tree leaf so it can serve as the default mutation-calling reference in Nextclade Web. The 498-tip tree spans 8 subgenogroups (A-H), and pathogen.json carries 3885 nucleotide mutation labels.

The PR description contradicts the shipped design: it says the dataset is "rooted on the Kowalik reference sequence" and adds "the Kowalik reference genome sequence in reference.fasta", but the file holds the inferred ancestor and the tree is rooted on that ancestor, exactly as the README's Scope section explains (F1). Its file list also omits tree.json, pathogen.json, sequences.fasta, and the collection registration, which carry the clade scheme, QC configuration, and examples.

Background

Coxsackievirus A10: taxonomy, genome, and the A-H genotype scheme [click to expand]

Taxonomy and reference strain

Coxsackievirus A10 is a serotype within the species Enterovirus alphacoxsackie (the ICTV binomial adopted in the 2023 Master Species List; formerly Enterovirus A), genus Enterovirus, family Picornaviridae, and is a positive-sense single-stranded RNA virus, Baltimore group IV (NCBI Taxonomy 42769). The reference is the Kowalik prototype, AY421767.1, 7409 nt, sequenced in the complete-genome survey of the species by Oberste et al., J Gen Virol 2004. There is no RefSeq genome for CVA10.

Genome organization

The genome is a single ~7.4 kb ORF: 5'UTR 1..744 and CDS 745..7326. The polyprotein is cleaved into capsid proteins (genome order VP4, VP2, VP3, VP1) and non-structural proteins 2A-2C and 3A, 3B (VPg), 3C (protease), 3D (RdRp). The dataset annotates all eleven mature peptides as separate CDS features, the convention across the ENPEN enterovirus datasets, so Nextclade reports amino-acid changes per mature protein. VP1 is the standard enterovirus typing target because its sequence correlates with serotype (Oberste et al., J Virol 1999) and lies in a recombination coldspot; partial VP1 sequencing is the routine surveillance assay, which is why a whole-genome dataset must handle VP1-only input.

Clade and lineage classification

Ji et al., Sci Rep 2018 defines genotypes A-G on the 894 nt VP1 gene with a 14.97% between-genotype divergence threshold, splits C into C1/C2, and reports H and I as candidates detectable only on partial VP1. Yang et al., J Med Virol 2025 works with the eight-genotype A-H scheme this dataset implements. Geographically, A is the prototype alone, B was a ceased Chinese lineage, C (especially C2) dominates Asia, D is the European lineage (also entering China after 2016), E is West/Central African, F is Indian, and G was first described from Taiwan. Clinically CVA10 is an established cause of hand-foot-and-mouth disease, herpangina, and onychomadesis, and has become a primary HFMD agent in parts of China as EV-A71 declined. Recombination is frequent in the non-structural P2/P3 regions while VP1 stays clonal (Wang et al., Virol Sin 2022), so a whole-genome placement can disagree with VP1 typing for a recombinant strain -- the reason VP1 remains the typing axis.

Non-blocking issues

🟡 F1. The reference role is described inconsistently across the README and PR [click to expand]

The dataset uses two sequences for two purposes: the inferred ancestor aligns and roots the tree, and Kowalik is the default reference for mutation calling in Nextclade Web. The configuration states this unambiguously (attributes["reference accession"] is ancestral_sequence; tree.json ref_nodes.default resolves to the AY421767 node). The prose disagrees with itself:

Location Says the reference is Correct?
README.md#L7 (table row) AY421767.1 No
README.md#L14 (Scope) ancestor aligns, Kowalik for mutation calling Yes
README.md#L52 (Reference types) ancestor is the mutation-calling default No
PR description dataset rooted on Kowalik No

The Reference-types sentence is inherited verbatim from ev-a71/README.md, so the same error is already live in that released dataset; the table row is a cva10-specific divergence, since cva16 and ev-a71 both put Static Inferred Ancestor in that row. A user reading the table concludes alignment is against the Kowalik GenBank sequence; a user reading the Reference-types section concludes reported mutations are relative to the ancestor. Neither matches what Nextclade does.

Suggestions:

  • Set the table row to Static Inferred Ancestor, matching the siblings, and rewrite the Reference-types bullet so the ancestor is the alignment reference and tree root and Kowalik is the mutation-calling default
  • Correct the PR description, and optionally fix the inherited sentence in ev-a71
🟡 F2. The clade scheme is uncited and the documented resolution is imprecise [click to expand]

The README advertises "subgenotype assignment" (README.md#L16) and states that naming "follows conventions established in published studies" without naming any. Three problems follow:

  • The A-H scheme has a citable source and none is given. Ji et al., Sci Rep 2018 defines genotypes A-G on the complete 894 nt VP1 gene with a 14.97% between-genotype divergence threshold and reports H (and I) as candidate genotypes detectable only on partial VP1 -- which is exactly the dataset's "H (VP1 only)" label. Yang et al., J Med Virol 2025 works with the full eight-genotype A-H scheme.
  • The delivered resolution is genotype level (A-H). Ji et al. further split genotype C into C1 and C2, so "subgenotype assignment" overstates what the tree returns; a C1 or C2 query receives only C.
  • The clade-founder example names B5, C4 (README.md#L56), which are EV-A71 subgenotypes copied from ev-a71/README.md; they do not exist in this tree. cva16 localized the same line to its own clades.

Suggestions:

  • Add a ## Citation section (the ev-d68 sibling already has one) citing Ji et al. 2018 and Yang et al. 2025
  • Describe the capability as genotype/genogroup assignment at A-H resolution, or add the C1/C2 subdivisions if intended
  • Replace B5, C4 with clades present here, for example C, F
🟡 F3. Two user-facing links are broken or non-resolving [click to expand]

Both links are rendered in the Nextclade Web dataset page:

  • The author-row link for Nadia Neuner-Jehle, https://eve-lab.org/people/nadia-neuner, returns HTTP 404 (README.md#L5). The working form, https://eve-lab.org/people/nadia-neuner-jehle/ (HTTP 200), is already used in tree.json meta.maintainers.
  • The ancestral_sequence tree leaf sets both accession and url to ancestral_sequence, producing the link https://www.ncbi.nlm.nih.gov/nuccore/ancestral_sequence, which resolves to no record because the inferred ancestor is not a GenBank sequence.

Suggestions:

  • Use https://eve-lab.org/people/nadia-neuner-jehle/ in the author row
  • Omit the url/accession for the ancestral tip, or point it at the upstream inferred-root.fasta
🟡 F4. Sparse clade G coverage makes genuine divergent G sequences fail QC [click to expand]

Clade G has only 5 tips, and the two clade-G example sequences that are not already tree leaves both come back bad:

Sequence Isolate Private subs Reversions QC
PQ889343 CMR-ENV02-N18 (Cameroon) 426 93 bad
PQ889344 CMR-ENV09-N22 (Cameroon) 432 94 bad

The clade call itself is correct; the tree simply holds almost no African or environmental G diversity (G is 5 tips, and only 13 of 498 tips are African). Unlike the five deliberately hard cases, these two are not listed under the # Outliers heading of the upstream resources/include_examples.txt, so they arrived through ordinary sampling and fail. The privateMutations thresholds were already deliberately raised upstream to cutoff: 250, typical: 50 (commit 7e5aae13), which is reasonable, but 426 and 432 are far beyond that. The user-facing effect is that genuine CVA10 sequences from under-sampled regions read as "suspect" when the real cause is thin reference coverage -- and environmental and African surveillance sequences are exactly what ENPEN members are likely to test.

Suggestions:

  • Add clade-G representatives spanning African and environmental diversity (PQ889343 and PQ889344 among them) so the clade has a reachable relative
  • If those two are deliberately excluded from the tree, list them under # Outliers in include_examples.txt so their QC status reads as intentional
🟡 F5. Unsequenced regions of the VP1-only clades appear as deletions in the reference tree [click to expand]

Clades E and H are seeded from partial VP1 records. In tree.json, their non-VP1 genome is reconstructed as nucleotide deletions rather than missing data: the E founder (NODE_0000125) carries 4832 deletions and the H founder (NODE_0000074) 5452, almost all outside VP1 (positions 2437-3330), and a shared internal branch (NODE_0000072) carries a further 1215. Read literally, these display as genome-wide deletions on lineages that were only ever partially sequenced.

The impact is limited to tree display, not analysis. A live CLI run confirms queries are unaffected: every example returns 0-9 total deletions, and the E-assigned example returns 0. Nextclade computes query mutations against reference.fasta, not against the tree's internal reconstruction, so the deletions do not propagate to placement or QC output. The upstream overhang-handling step explains the representation.

Suggestions:

  • In the tree build, encode uncovered regions of partial records as missing (N) rather than as deletions, so the E and H branches read as VP1-only
  • Substitutions that observed sequence supports should be kept (see N1)
🔵 F6. Annotation sequence identifiers disagree [click to expand]

One reference sequence carries three identifiers: reference.fasta header ancestral_sequence, GFF3 seqid AY421767.1, and tree.json meta.genome_annotations[*].seqid resources/reference.gbk (a build-local path). Numeric coordinates agree and Nextclade runs cleanly (0 warnings on the reference), and the same pattern exists in the released siblings, so nothing breaks; but external tools that join FASTA and annotation by identifier cannot, and the tree exposes a path that is not a sequence.

Suggestions:

  • Use the reference identifier consistently (AY421767.1 in the tree seqid to match the GFF3), keeping Kowalik as documented coordinate provenance
🔵 F7. Collection Date is a 261-value categorical coloring [click to expand]

tree.json declares { "key": "date", "type": "categorical" } with a 261-entry scale, and no tip carries num_date (dates are strings like "2008-XX-XX"). Coloring by date therefore produces 261 unordered categories, so adjacent years get unrelated colors. The ev-d68 sibling uses "type": "ordinal" for the same field and reads correctly; cva16 and ev-a71 share the categorical form, so this is a collection-wide pattern.

Suggestions:

  • Change the coloring type to ordinal, or emit num_date on tips with a resolvable date
🔵 F8. CHANGELOG grammar [click to expand]

CHANGELOG.md#L3 reads "Initial release of an Coxsackievirus A10 dataset". The changelog is shown in Nextclade Web.

Suggestions:

  • Change an to a
🔵 F9. tree.json has no trailing newline [click to expand]

Both the source and generated tree.json end directly after }; every other text file in the dataset ends with a newline. The rebuild script minifies the tree without a trailing newline, so a source-only fix would not persist -- the change belongs in the upstream JSON writer. Cosmetic (noisier diffs on future updates).

Suggestions:

  • Emit a trailing newline when the upstream workflow writes the tree

Clade distribution

8 subgenogroups plus the inferred ancestor, 498 tips [click to expand]
Clade Tips %
C 361 72.5
F 70 14.1
D 49 9.8
E (VP1 only) 8 1.6
G 5 1.0
H (VP1 only) 2 0.4
A 1 0.2
B 1 0.2
unassigned (inferred ancestor) 1 0.2
Total 498

Sampling is heavily weighted toward China (334/498) and toward 2018-2019, thinning after 2022; dated tips span 2008-2025. The balance is defensible against the literature:

  • C and D dominance: genotype C dominates Asian circulation and D is the European lineage (Ji et al. 2018; Wang et al., Virol Sin 2022), so the 361/49 split is expected
  • Single-tip clades A and B: genotype A is prototype-only (the one A tip is Kowalik itself) and B ceased circulating around 2009, so single tips are correct
  • VP1-only clades E and H: both rest on partial VP1 data, so their sparse whole-genome representation is expected

The real gaps are clade G (F4) and recent European D diversity: the earliest D tip is 2014, so the outbreak lineages that defined D are absent.

Validation summary

Structural checks and official Docker run [click to expand]

GFF3 annotation

  • 11/11 CDS lengths divisible by 3; the eleven mature peptides form a continuous polyprotein
  • Polyprotein starts ATG at 745 and is followed by TAA at 7324-7326, consistent with GenBank AY421767.1 CDS 745..7326; 0 internal premature stops
  • defaultCds is VP1, which exists in the GFF3
  • Coordinates are 1-based and identical to meta.genome_annotations in tree.json

Reference

  • 7409 nt, 0 ambiguous bases; full polyprotein translates to 2194 aa with 0 premature stops
  • Header ancestral_sequence matches attributes["reference accession"]; reference name is distinct and informative

Tree

  • Valid Auspice v2, 498 tips; both ancestral_sequence (alignment reference) and AY421767 (mutation-calling reference) are present as leaves, so ref_nodes resolves
  • root_sequence absent, consistent with all three siblings and not required when reference.fasta is supplied
  • Dated tips span 2008 to 2025-07-11
  • Fail: seqid is a build-local path (F6); date coloring is categorical over 261 values (F7)

Generated output and registration

  • README, CHANGELOG, GFF3, reference.fasta, sequences.fasta byte-identical between source and data_output/
  • pathogen.json differs only by the injected version block (tag: unreleased); tree.json differs only in serialization
  • Registered in data/enpen/collection.json and data_output/index.json; minimizer index includes cva10

Nextclade CLI run (nextstrain/nextclade, digest sha256:a7b995e2, Nextclade 3.23.0)

  • Reference: QC good, 0 substitutions, 0 private mutations, no warnings
  • 30 examples: 24 good, 2 mediocre, 4 bad, with no frameshifts, no premature stops, and no errors; per-query deletions stay within 0-9
  • All 4 bad and both mediocre are driven by privateMutations; the 8 "empty gene/all gaps" warnings come from partial or VP1-only inputs (expected)
  • Five of the six non-good examples are deliberate upstream outliers; only the two clade-G Cameroon sequences are unaccounted for (F4)

Notes

Click to expand
  • N1. The VP1-only handling of E and H is correct, not a defect. An earlier concern that E/H carry substitutions and mutation labels outside VP1 does not hold: the E reference records extend well beyond VP1 (JN255588.1 is 1422 nt, JX307651.1 is 1399 nt against an 894 nt VP1), so those substitutions are supported by observed sequence. Every questioned label is also co-assigned to other clades (2205C to C/F/G/E, 3333G to D/H/G/B), and mutation labels annotate observed mutations for QC weighting -- they do not drive clade assignment, which is phylogenetic placement. Naming the limitation in the E (VP1 only) / H (VP1 only) labels means users see it in the results table.
  • N2. The Static Inferred Ancestor design is sound and consistent with the released cva16 and ev-a71 siblings. Kowalik sits ~22-25% divergent in VP1 from circulating genotypes, making it a poor alignment anchor; keeping its coordinate system preserves comparability with the literature while retaining it as a leaf keeps it available for mutation calling.
  • N3. The reference accession is correct on the science: no RefSeq (NC_*) genome exists for CVA10, and AY421767.1 is the community prototype from Oberste et al., J Gen Virol 2004.
  • N4. Clade assignment is robust to sequence length: re-running examples with only the VP1 region reproduces the whole-genome clade call, consistent with VP1 being a recombination coldspot and with the VP1-centred defaultCds.
  • N5. The strict minSeedCover of 0.84 (the highest in the collection) acts as a cross-serotype filter: CVA16 whole genomes are rejected or flagged bad, which makes the README's "Related Enteroviruses" warning empirically accurate.
  • N6. The privateMutations cutoff of 250 is deliberate and documented upstream (commit 7e5aae13), not a copied default; the median private-mutation count across runtime-placed examples is about 42, so the threshold is well placed for typical input.
  • N7. The path enpen/enterovirus/cva10 is consistent with the released siblings and correctly registered; the mixed separator style across the collection follows the official abbreviations and is already locked by released datasets.
  • N8. data_output/index.json advertises "clades": 9, one more than the eight subgenogroups, because unassigned on the ancestral tip is a distinct clade_membership value. This count comes from scripts/rebuild (tree_find_clades), which counts distinct values across the whole tree, so it is repo tooling behaviour affecting the collection (cva16 and ev-a71 count the same way), not something the dataset author sets or can change in this PR. Informational for repo maintainers.

This branch was successfully deployed

1 active deployment
refs/pull/466/merge — fefe1a7d Deployed Sep 15, 2026 by ivan-aksamentov via build-and-deploy-datasets #1524
refs/heads/add-cva10 — fefe1a7d Deployed Sep 15, 2026 by ivan-aksamentov via build-and-deploy-datasets #1523
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants