Exercises
Write queries against the open SARS-CoV-2 dataset. Results are automatically compared with the reference answer. Each exercise includes an explanation and links to the relevant documentation.
●
Fixed training environment
The exercises run against the SARS-CoV-2 dataset preprocessed by Nextstrain using data from INSDC and RKI and hosted by CoV-Spectrum.
https://gs-staging-1.int.genspectrum.org/open/v2/silo
- Exercise 1Count sequences from SwitzerlandCount all sequences from Switzerland.
- Exercise 2Retrieve Basel sequencesRetrieve the 20 most recent sequences from Basel-Stadt, Switzerland, showing the GenBank accession, date, unaligned nucleotide sequence, aligned nucleotide sequence and S amino acid sequence.
- Exercise 3Lineage counts in early 2021For sequences collected between 1 January 2021 and 30 June 2021, count the sequences per pangoLineage and return the 20 most frequent lineages, most frequent first.
- Exercise 4Compare worldwide and Spanish lineage countsIdentify the most common 50 pango lineages worldwide and return their worldwide sequence count and their sequence count from Spain. Sort by the worldwide count, highest first.
- Exercise 5Submissions by ISO weekSummarize sequences collected in 2024 by ISO calendar week. Return the ISO week and sequence count, ordered chronologically.
- Exercise 6Sequences with a mutationFor sequences that carry a mutation at nucleotide position 23403, show the strain, country, date and pangoLineage. Order by strain and return the first 20 rows.
- Exercise 7Mutations within a lineageList the nucleotide mutations (on the 'main' sequence) that occur in at least 5% of sequences belonging to lineage B.1.1.7 including its sublineages. Return at most 20 rows.
- Exercise 8Combine multiple conditionsCount the sequences from Germany that belong to B.1.1.7 (including sublineages) and carry at least 2 of these three nucleotide mutations: T at position 241, T at position 3037, G at position 23403.
- Exercise 9Paginated resultsOrder the sequences by strain, skip the first 50 and return the next 25, showing only the strain, country and date.
- Exercise 10Amino acid insertionsList the 20 most common amino acid insertions in the S protein. For each insertion, show its position, inserted symbols and the number of sequences carrying it, with the most common first.
- Exercise 11Filter with a set of valuesCount the sequences from Germany, France and Italy, broken down by country, most frequent first.
- Exercise 12Regex filter and a computed columnFor sequences whose division matches the regular expression 'Basel.*', return the strain together with an added column 'area' set to the constant 'Basel'. Order by strain and return the first 10 rows.
- Exercise 13Amino acid mutations grouped by targetAmong sequences from Switzerland, take the amino acid mutations on the S gene that occur in at least 10% of sequences, then count how many of those mutations lead to each resulting symbol (mutationTo). Most frequent first.
- Exercise 14Mutation profile distanceCount the sequences whose S gene is within 2 amino acid differences of a profile that has Y at position 501 and R at position 452.
- Exercise 15Co-occurring S protein changesIn the sequences from Switzerland, investigate the co-occurrence patterns at S protein positions 69, 70 and 501. Return one row for each combination together with the count. Order the combinations by count descending so the most common S-position pattern appears first.
- Exercise 16Compare recent German and US submissionsBuild one harmonized table for recent SARS-CoV-2 submissions from Germany and the USA. Include the strain name, collection date, pango lineage and a place column.
- Exercise 17Find Swiss lineages absent from ArgentinaWhich Pango lineages were observed in Switzerland but never in Argentina? Return the 20 lineages with the largest Swiss sequence count, highest first.
