Repository navigation
Conversation
|
Note ✨ Claude Opus 5.5 (with help from GPT-5.6 Sol) is responding on behalf of Ivan First review of PR #479, adding the TestingObservedWhat the dataset is, and where it diverges from its description [click to expand]The PR adds a whole-genome HCV dataset for all genotypes (1-8). The dataset has these properties:
The dataset is built by the new nextstrain/hcv workflow. The PR body contains only a preview link. The four source commits are the initial import, a GFF phase-column fix, Discrepancies between the description and the data:
BackgroundHCV taxonomy, H77 reference, genome organization, genotype classification [click to expand]Taxonomy and reference strainHCV is a positive-sense single-stranded RNA virus (Baltimore group IV). ICTV MSL #41 (ratified 2026) renamed the species to Orthohepacivirus hominis, in genus Orthohepacivirus and the new family Hepaciviridae (ICTV taxon history, Orthohepacivirus). The previous name was Hepacivirus hominis in Flaviviridae (MSL #38, Postler et al., Arch Virol 2023). H77 comes from a US patient who was infected in 1977 (Ogata et al., PNAS 1991). Two different infectious cDNA clones of H77 are in GenBank:
Genome organizationThe 9.6 kb genome encodes one polyprotein that is cleaved into core, E1, E2, p7, NS2, NS3, NS4A, NS4B, NS5A and NS5B. The dataset annotates these ten mature peptides. Only core starts with ATG, and NS5B is followed by the polyprotein stop codon (TGA), which is expected for mature peptides.
Genotype / subtype classificationICTV currently recognizes 8 genotypes and 93 confirmed subtypes. Table 1, dated January 2026, gives one reference genome per subtype. ICTV applies these rules (ICTV Hepacivirus; Smith et al. 2014):
Genotype 8 was described in 2018 from four patients from Punjab, India (Borgia et al., J Infect Dis 2018). Only subtype 8a is confirmed. Further genotype 8 genomes from the UK are in GenBank (PP092205, submitter title "two new subtypes of hepatitis C virus genotype 8"). Epidemiology and why subtype mattersWHO estimates 47 million people with chronic HCV in 2024 (WHO fact sheet). Genotype 1 is about 44-46% of infections and genotype 3 is 25-30% (Messina et al., Hepatology 2015; Polaris Observatory 2017). With pan-genotypic antivirals, subtype still matters. Some "unusual" subtypes (1l, 3b, 3g, 4r, 6u, 6v) carry NS5A polymorphisms that reduce susceptibility to direct-acting antivirals (Nguyen et al., J Hepatol 2020). Unusual genotype 1 subtypes are over-represented among treatment failures (Vo-Quang et al., Hepatology 2023). A wrong subtype call is therefore more harmful than no call (F2). These tools also assign HCV genotype and subtype: Blocking issues🔴 F1. Dataset path and CHANGELOG name AF009606, but the reference is NC_038882.1 (= AF011751) [click to expand]The dataset path
NCBI states that Dataset paths are immutable after release (curation guide). If this ships, the identifier permanently names an accession that the dataset does not use. The dataset list would show Suggestions (pick one):
🔴 F2. Subtypes missing from the tree are often called as a wrong subtype [click to expand]The tree contains 46 of the 93 confirmed subtypes. The README says that sequences of unrepresented subtypes "are assigned the genotype, but may not get a subtype assignment". To test this, I ran the reference genome of every confirmed subtype from ICTV Table 1 that NCBI returned (92 genomes) through the dataset with Nextclade 3.23.0:
The genotype is always right. The subtype is wrong for 12 of the 60 genomes that are not in the tree:
The tree labels also disagree with ICTV for three reference tips:
The cause is that the subtype of the nearest labeled clade is inherited, and clade labels come from NCBI metadata. The PR does not change the Suggestions:
🔴 F3. A genotype 8 genome gets no genotype [click to expand]The README says the dataset covers "all known genotypes (1-8)". Genotype 8 has one tip (
Other public genotype 8 genomes are not in the tree, because NCBI has no genotype field for them and the workflow annotates only Genotype 7 has a similar weakness. It has 3 tips, but Suggestions:
Non-blocking issues🟡 F4. Placement masks are in pathogen.json, where Nextclade does not read them [click to expand]
I moved the same five ranges into a copy of Suggestions:
{
"meta": {
"extensions": {
"nextclade": {
"placement_mask_ranges": [
{ "begin": 2045, "end": 2070 },
{ "begin": 7035, "end": 7060 },
{ "begin": 7380, "end": 7490 },
{ "begin": 7560, "end": 7590 },
{ "begin": 9375, "end": 9599 }
]
}
}
}
}
🟡 F5. Genotype 1 is missing from the dataset-suggestion fingerprint [click to expand]
Genotype 1 is the most common genotype (see Background). Nextclade Web does not suggest this dataset for most genotype 1 partial sequences. Suggestions:
🟡 F6. NS5A V3 alignment artifacts trigger frameshift QC [click to expand]Frameshift QC is "mediocre" for 5/20 example sequences (6a x3, 2b x2) and for 2/7 additional references (3a, 7). All flagged frameshifts are in NS5A codons 376-408, which is variable region V3 (see Background). Each one is a compensated indel pair, for example:
These are alignment artifacts against a genotype 1a reference, not real frameshifts. Amino acid changes that Nextclade reports in this NS5A window for non-1 genotypes are unreliable. None of these changes removed the artifacts:
Suggestions:
🟡 F7. README has unfilled TODO placeholders [click to expand]Nextclade Web shows the README to users. It contains three placeholders:
Suggestions:
🟡 F8. Core/E1 typing fragments often fail to align [click to expand]I cut fragments from the 27 examples and additional references (reference coordinates from the Nextclade alignment) and ran them through the dataset:
Core/E1 is one of the two standard typing regions (Murphy et al. 2007). For 9 genomes (2b x3, 3a x4, 5a, 6t), the fragment fails with "seed alignment was unable to find any matches". Suggestions:
🔵 F9. CHANGELOG has no trailing newline [click to expand]
Suggestions:
🔵 F10. No compatibility block [click to expand]
Suggestions:
🔵 F11. The tree has duplicate RefSeq/GenBank tips [click to expand]Five tip pairs are the same record under a RefSeq and a GenBank accession (zero-length sister branches):
These pairs inflate the tip counts. For genotype 7, 3 tips are only 2 distinct genomes. Suggestions:
🔵 F12. Tree defines a coloring with no data [click to expand]
Suggestions:
Clade distribution8 genotypes, 46 subtypes, 182 tips [click to expand]
Validation summaryValidation checks [click to expand]Reference
GFF3 annotation
pathogen.json
Tree
Examples
Generated output
Nextclade CLI run (3.23.0, docker)
NotesClick to expand
|
|
Probably too early, but I had some spare tokens before reset, so decided to run my AI review |
Preview link: https://master.clades.nextstrain.org/?dataset-server=gh:@hcv&dataset-name=nextstrain/hcv/AF009606