This filter removes outlier markers with too many SNP number per locus/read. The data requires snp and locus information (e.g. from a VCF file). Having a higher than "normal" SNP number is usually the results of assembly artifacts or bad assembly parameters. This filter is population-agnostic, but still requires a strata file if a vcf file is used as input.
Filter target: Markers.
Statistics: The number of SNPs per locus.
For reference-based VCFs, this statistic is the number of VCF records sharing
the same stored LOCUS. If the VCF does not preserve a RAD/read locus
identifier, each record has a unique LOCUS; the function then reports
one SNP per locus and leaves the data unchanged. The nucleotide length of a
FreeBayes REF allele is not interpreted as a SNP count because complex
substitutions and indels can span several bases without representing several
independently stored SNP records.
Usage
filter_snp_number(
data,
strata = NULL,
interactive.filter = TRUE,
filter.snp.number = NULL,
filename = NULL,
parallel.core = parallel::detectCores() - 1,
verbose = TRUE,
...
)Arguments
- data
A tidy genomic data frame or another genomic object supported by the calling function.
- strata
Optional strata file or object containing individual and group information. Default:
strata = NULL.- interactive.filter
Logical indicating whether an interactive filtering session may display diagnostics and ask for thresholds. Default:
interactive.filter = TRUE.- filter.snp.number
(integer) This is best decided after viewing the figures. If the argument is set to 2, locus with 3 and more SNPs will be blacklisted. Default:
filter.snp.number = NULL.- filename
(optional) Name of the filtered tidy data frame file written to the working directory (ending with
.tsv) Default:filename = NULL.- parallel.core
Number of workers available for parallel operations. Default:
parallel.core = parallel::detectCores() - 1.- verbose
Logical. Display progress messages. Default:
verbose = TRUE.- ...
Additional arguments passed to lower-level screening or filtering functions.
Value
The filtered data in the same representation as the input. GDS marker metadata and active variants are updated in place. Diagnostic files, marker lists, and filtering parameters are written to the output folder.
STACKS loci and FreeBayes haplotypes
STACKS reconstructs RAD loci and preserves a locus identifier shared by the SNPs observed on the same read or locus. In that setting, asking whether a roughly 70-100-bp sequence contains an unusually large number of SNPs is meaningful: dense polymorphism can be consistent with paralogy, poor locus assembly, misalignment, or unexpectedly high local diversity.
FreeBayes is haplotype-aware in a different sense. During variant calling it
evaluates nearby alleles jointly and may emit a SNP, multi-nucleotide
polymorphism, indel, or complex allele as one VCF record. This does not imply
that the record represents one reconstructed RAD locus, and the number of
bases in its REF or ALT allele is not the number of SNPs on a read. Unless an
upstream shared RAD/read locus identifier was retained, genomic coordinates
cannot unambiguously recover the original read boundaries. Consequently,
filter_snp_number() does not infer them and skips filtering when every
VCF record has a unique LOCUS.
For reference-based FreeBayes data without retained locus identifiers, use genotype quality, allele balance, depth and coverage, excess heterozygosity, mapping diagnostics, local variant density, and genomic LD as complementary diagnostics. Local variant-density windows must be interpreted as genomic windows, not automatically as individual RAD reads.
Interactive version
The function first displays and writes the distribution of SNPs per locus and the effect of candidate thresholds. It then asks:
"Do you still want to blacklist markers? (y/n):"If yes, choose
1to use the upper boxplot-outlier statistic or2to enter a threshold.With option 2, answer
"Enter the maximum number of SNP per locus allowed:".
All SNPs in loci containing more than the selected number are blacklisted.
Answering no leaves the data unchanged. Use
interactive.filter = FALSE and provide filter.snp.number
explicitly for a reproducible analysis.
Examples
if (FALSE) { # \dontrun{
genome <- genometranslator::read_genome(
data = "turtle.vcf",
strata = "turtle.strata.tsv"
)
# Inspect the SNP-per-locus distribution interactively.
genome <- radr::filter_snp_number(data = genome)
# Alternatively, use a separate unfiltered GDS for a scripted run that
# retains loci containing at most four SNPs.
scripted_genome <- genometranslator::read_genome("turtle_scripted.gds")
scripted_genome <- radr::filter_snp_number(
data = scripted_genome,
interactive.filter = FALSE,
filter.snp.number = 4
)
} # }
