| Version: | 26.3 |
| Date: | 2026-09-05 |
| Title: | Simulation and Data Analysis for Plant Breeders |
| Description: | Provides tools for simulation of plant breeding programs as described, for example, by Melchinger and Frisch (2023) <doi:10.1007/s00122-023-04446-3>, prediction of segregation variance (Osthushenrich, Frisch and Herzog (2017) <doi:10.1371/journal.pone.0188839>), genomic prediction (Hofheinz and Frisch (2014) <doi:10.1534/g3.113.010025>), linkage disequilibrium based haplotype construction, and planning of marker assisted back crossing programs. |
| Suggests: | knitr, rmarkdown |
| VignetteBuilder: | knitr |
| Depends: | R (≥ 4.0.0) |
| License: | CC0 |
| NeedsCompilation: | yes |
| Packaged: | 2026-09-05 11:22:03 UTC; frisch-m |
| Author: | Matthias Frisch [aut, cre], Hans Peter Maurer [ctb], Philipp Heilmann [ctb] |
| Maintainer: | Matthias Frisch <matthias.frisch@uni-giessen.de> |
| Repository: | CRAN |
| Date/Publication: | 2026-09-15 10:40:21 UTC |
Maximum Allele Code Used by SelectionTools
Description
Internal constant giving the largest biological allele code supported by SelectionTools.
Usage
.GS_ALLELE_MAX
Format
An integer scalar with value 254.
Details
The value is used when validating allele codes in the R interface. The separate internal missing-value code is not a biological allele code. This constant is intended for internal use and should not be modified.
Value
An integer scalar with value 254, giving the largest biological allele
code supported by the SelectionTools R interface. The separate missing-value
code is not included in this maximum.
Maximum Number of ALT Alleles in STvcf Data
Description
Internal constant giving the maximum number of ALT alleles supported for one marker in an STvcf object.
Usage
.GS_STVCF_MAX_ALT
Format
An integer scalar with value 253.
Details
Together with the REF allele, at most 254 biological allele states are supported for one marker. Missing data use a separate internal state. This constant is intended for internal use and should not be modified.
Value
An integer scalar with value 253, giving the maximum number of ALT
alleles supported for one marker in an STvcf object. Together with the REF
allele this permits 254 biological allele states.
Technow Hybrid Phenotype Data
Description
Phenotypic data for a maize hybrid data set used for examples of genomic prediction in SelectionTools. The records identify dent and flint parents, their hybrid, and grain-yield performance.
Usage
data(DT_technow)
Format
A data frame containing the variables dent, flint,
hy, and GY. The first two variables identify the parental
lines, hy identifies the hybrid, and GY contains grain-yield
phenotypes.
Source
Data supplied with SelectionTools for genomic-prediction examples.
Technow Dent Marker Data
Description
Marker data for the dent parental lines corresponding to the hybrid
phenotype data in DT_technow. The data are used to construct
genomic relationship matrices and for genomic-prediction examples.
Usage
data(Md_technow)
Format
A marker-data matrix or data frame with dent parental lines in rows and markers in columns. Row names identify the parental lines and the entries contain marker genotypes.
Source
Data supplied with SelectionTools for genomic-prediction examples.
Technow Flint Marker Data
Description
Marker data for the flint parental lines corresponding to the hybrid
phenotype data in DT_technow. The data are used to construct
genomic relationship matrices and for genomic-prediction examples.
Usage
data(Mf_technow)
Format
A marker-data matrix or data frame with flint parental lines in rows and markers in columns. Row names identify the parental lines and the entries contain marker genotypes.
Source
Data supplied with SelectionTools for genomic-prediction examples.
append.population()
Description
Appends a copy of one population to another.
Usage
append.population(NameP1, NameP2)
Arguments
NameP1 |
Population which takes up the copied individuals |
NameP2 |
The population to be appended |
Details
The individuals of NameP2 are copied to the end of NameP1; NameP2 is not modified. Raw chromosome data are copied for every appended individual. Cached genotype and genetic-value categories are retained only when they can remain complete and internally consistent for all active individuals of the destination; otherwise the corresponding derived category is cleared.
Value
Returns the same invisible implementation-level list as population.append(). The function is a compatibility wrapper and the meaningful result is the delegated population operation.
Switch SelectionTools output to a quiet mode
Description
Switch SelectionTools output to a quiet mode.
Usage
be.quiet()
Value
Returns the same invisible implementation-level list as set.info.level(0); the main effect is switching package information output to quiet mode.
concat.population()
Description
Concatenates two populations and removes the second population.
Usage
concat.population(NameP1, NameP2)
Arguments
NameP1 |
Recipient population |
NameP2 |
Name of the population to be concatenated |
Details
The individuals of NameP2 are appended to NameP1. After successful concatenation, NameP2 is removed. Complete individual records are moved or copied as required so chromosome data and any valid per-individual derived information remain associated with the correct individuals.
Value
Returns the same invisible implementation-level list as population.concat(). The function is a compatibility wrapper and the meaningful result is the delegated population operation.
copy.population()
Description
Copies all or a contiguous part of a simulation population into a destination population.
Usage
copy.population(NameP1, NameP2, start=1, n=-1)
Arguments
NameP1 |
Name of the population to be created |
NameP2 |
Name of the original population |
start |
Index of the first copied individual |
n |
Number of individuals to be copied. Default: n=-1 means all individuals |
Details
Copying starts at the one-based position start. A negative n requests the remaining individuals through the end of the source; a positive value requests that many individuals subject to the source bounds. The source population is not modified. Raw chromosome data are copied, and genotype or genetic-value data are copied only when the corresponding source category is complete for all active source individuals.
Value
Returns the same invisible implementation-level list as population.copy(). The function is a compatibility wrapper and the meaningful result is the delegated population operation.
cross()
Description
Generates a diploid progeny population by crossing individuals drawn from two parental populations.
Usage
cross(NamePg, NameP1, NameP2, NoPg = -1, self = 0, mode = 1,
crossing.scheme = 0)
Arguments
NamePg |
Character string naming the progeny population. |
NameP1 |
Character string naming parental population 1. |
NameP2 |
Character string naming parental population 2. |
NoPg |
Number of progeny to generate. A positive value creates that many individuals; |
self |
Numeric selfing rate passed to the crossing routine. |
mode |
Integer mating-mode flag controlling self-incompatibility behavior. |
crossing.scheme |
Integer 0 to 3 selecting random or deterministic parent choice in the two parental populations. |
Details
The simulation genome must be diploid. The progeny population must have a name different from both parental populations. Existing contents under the progeny name are replaced when a new progeny population is constructed.
For positive NoPg, exactly that many progeny individuals are
generated. The special value -1 uses the already allocated size
of an existing progeny population; it is therefore meaningful only when
such a destination population already exists with positive allocation.
The four crossing schemes control whether the parent chosen from each
parental population is random or deterministic. Deterministic selection
cycles through the active individuals of the corresponding parent
population. self controls selfing, and mode controls
whether the mating logic applies self-incompatibility.
Each progeny receives newly simulated chromosome segments produced by meiosis. Cached marker genotypes and evaluated genetic values are not inherited as authoritative derived data; they can be generated from the new chromosomes when needed. If chromosome construction fails, the incomplete progeny population is removed rather than left partially generated.
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of creating the requested crosses.
Note
Crossing scheme 0 selects both parents randomly; scheme 1 selects the first parent randomly and the second deterministically; scheme 2 does the reverse; scheme 3 selects both deterministically.
Return or inspect parameters associated with a data object
Description
Return or inspect parameters associated with a data object.
Usage
data.params(dta)
Arguments
dta |
Input data object. |
Value
A named list describing the supplied data object, with components:
colhead |
Parsed column headings identifying populations and individuals. |
rowhead |
Parsed row headings identifying markers and alleles. |
no.mar |
Number of markers. |
no.all |
Number of alleles per marker. |
no.pop |
Number of populations. |
no.ind |
Numbers of individuals per population. |
ind.list |
Individual numbers within populations. |
mar.list |
Marker names. |
pop.list |
Population names. |
define.effects()
Description
Defines effects that can be used for selection
Usage
define.effects(name, description)
Arguments
name |
Name of the effect |
description |
Description of the effect |
Details
The syntax for the description string is
Mean, Classname Allele [Allele] Distribution Parameter [,...]
| Mean | The population mean |
| Classname | A class of Loci to which effects are assigned to |
| Allele | The allels to which the effects are assigned to |
| [Allele] | Optional. Is used to define dominance effects |
| [, ...] | An arbitrary number of effects can be defined separated by commas |
| Distribution | normal 0/0.01 or uniform 0 |
Value
An integer status returned by the final unlink() call that removes the temporary effect-definition file. The meaningful result is the side effect of defining/loading the requested effects.
define.effects.df()
Description
Automatically generates effects for loci: effect is always 1 for allele 1.
Usage
define.effects.df ( LocNames )
Arguments
LocNames |
Name of the loci for which standard effects of 1 are generated |
Details
Population specified via PopName is accessed in the background and must therefore be available in the backend.
Value
NULL. The function is called for its side effect of defining effect maps for the locus names in the supplied data frame; invalid pre-existing definitions return NULL invisibly.
Examples
map <- data.frame(chrom = 1:2,
pos = c(0.01,0.02),
name = c("form","color"))
define.genome(map)
define.effects.df(map$name)
Define a genome from a linkage map and optional chromosome lengths
Description
Define a genome from a linkage map and optional chromosome lengths.
Usage
define.genome(map, length.vec=NULL)
Arguments
map |
Linkage map or map object. |
length.vec |
Vector of chromosome lengths. If a chromosome length is zero, it is replaced by 0.01 and a warning is issued. |
Value
NULL on successful completion and on the documented validation exits. The function is called for its side effects of setting genome parameters and defining the linkage map.
define.map()
Description
Defines and installs a simulation linkage map from a compact textual map description.
Usage
define.map(description)
Arguments
description |
A string containing a description of the linkage map |
Details
The description is first converted to a temporary map file by generate.map.file and is then loaded through the simulation linkage-map loader. The operation therefore uses the same map validation and ordering rules as file-based simulation-map loading.
Description elements specify a locus name, class, and position rule. A position can identify a chromosome and map position directly or can request generated placement such as random, terminal, or nonterminal positions according to the existing description syntax. A multiplicity prefix can request several loci from one specification.
The function replaces the currently active simulation linkage-map definition only after its generated description has been accepted by the underlying map-generation and loading operations.
Value
An invisible list returned by generate.map.file() or linkage.map.load(). Its retval component is the status code of the map-generation/loading step; the main effect is defining the simulation map.
Issue information for a deprecated function or object name
Description
Issue information for a deprecated function or object name.
Usage
deprecated(NameOld, NameNew)
Arguments
NameOld |
Previous function or object name. |
NameNew |
Replacement function or object name. |
Value
NULL. This helper has no result object.
dh()
Description
Generates a diploid population of doubled-haploid lines from a parental population.
Usage
dh(NamePg, NameP1, NoPg=-1)
Arguments
NamePg |
Character string naming the doubled-haploid progeny population. |
NameP1 |
Character string naming the parental population. |
NoPg |
Number of progeny. A positive value creates that many lines; |
Details
For each progeny, a parental individual is chosen and a haploid meiotic product is generated. The selected haploid chromosome set is then doubled so that both homologues of each progeny carry the same recombined haplotype.
The simulation genome must be diploid. For positive NoPg, that many doubled-haploid progeny are generated. The special value -1 uses the allocation of an existing progeny population and therefore requires an existing destination with positive allocation.
The destination population is constructed from raw chromosome segments. Cached genotype and evaluation data can subsequently be recalculated. If meiosis or chromosome construction fails, the incomplete progeny population is removed.
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of creating doubled-haploid progeny.
divide.population()
Description
Splits individuals from one simulation population into a second population.
Usage
divide.population(NameP1, NameP2, NoI=-1)
Arguments
NameP1 |
Name of the population to be created |
NameP2 |
Name of the population to be divided |
NoI |
Number of individuals to be separated from population NameP2 |
Details
The first requested individuals are
separated from NameP2 and placed in NameP1; the
remaining individuals stay in NameP2. Complete individual
structures are moved, so chromosome segments, information fields,
genotype data, and evaluations that belong to an individual move
together.
Value
Returns the same invisible implementation-level list as population.divide(). The function is a compatibility wrapper and the meaningful result is the delegated population operation.
Set effect weights for all effects
Description
Set effect weights for all effects.
Usage
effect.weight.set.all(weight)
Arguments
weight |
Weight value. |
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of setting all effect weights.
Remove a stored effect map
Description
Remove a stored effect map.
Usage
effmap.remove(Name)
Arguments
Name |
Name identifying the object to be processed. |
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of removing the requested effect map.
Remove all stored effect maps
Description
Remove all stored effect maps.
Usage
effmap.remove.all()
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of removing all effect maps.
Evaluate all loci in a population
Description
Evaluate all loci in a population.
Usage
evaluate.all.loci(PopName,
loci = paste(as.vector(unlist(linkage.map.get()[,3])), collapse=" "))
Arguments
PopName |
Name of the population. |
loci |
Locus or loci to process. |
Value
The same data.frame structure as evaluate.locus(), evaluated for all loci in the current map: rows are loci and columns are homozygote, heterozygote, and missing.
evaluate.allele()
Description
Calculates allele frequencies
Usage
evaluate.allele(PopName, Loci, missing=0)
Arguments
PopName |
Name of the population to be evaluated |
Loci |
Names of the loci to be evaluated |
missing |
Function argument. |
Value
A data.frame containing one column for each evaluated locus plus a count column giving the number of observations for each allele combination. Returns NULL invisibly if no evaluable loci/result are available.
Note
Only positive alleles (without zero) are allowed!
Calculate or return allele frequencies for selected loci
Description
Calculate or return allele frequencies for selected loci.
Usage
evaluate.allele.freq(PopName, Loci, missing=0)
Arguments
PopName |
Name of the population. |
Loci |
Locus or loci to process. |
missing |
Value or code used for missing data. |
Value
A data.frame containing the allele counts returned by evaluate.allele() and an added freq column containing relative frequencies. Returns NULL invisibly when the request cannot be evaluated.
evaluate.allele.freq2()
Description
Calculates allele frequencies
Usage
evaluate.allele.freq2 ( PopName, Loci, missing = 0 )
Arguments
PopName |
Name of the population to be evaluated |
Loci |
A vector containing names of Loci |
missing |
set bg = 1 if genotypes containing any missing value (NA) should be dropped for generating the evaluation table |
Details
Population specified via PopName is accessed in the background and must therefore be available in the backend.
Value
A data.frame containing the allele counts returned by evaluate.allele2() and an added freq column containing relative frequencies. Returns NULL invisibly when the request cannot be evaluated.
Examples
map <- data.frame(chrom = 1:2,
pos = c(0.01,0.0),
name = c("form","color"),
class = paste0("trait",1:2))
define.genome(map)
init.population("P1",homozygote(1))
init.population("P2",homozygote(2))
cross("F1","P1","P2",10)
cross("F2","F1","F1",1000)
evaluate.genotype2("F2", map$name)
evaluate.genotype2("F2", map$name, alleles = c(1,1,2,2))
Evaluate allele information for selected loci in a population
Description
Evaluate allele information for selected loci in a population.
Usage
evaluate.allele2(PopName, Loci, missing=0)
Arguments
PopName |
Name of the population. |
Loci |
Locus or loci to process. |
missing |
Value or code used for missing data. |
Value
A data.frame containing one column for each evaluated locus plus a count column giving the number of observations for each allele combination. Returns NULL invisibly if no evaluable loci/result are available.
evaluate.genome()
Description
Computes the frequency distribution of an allele across the whole genome.
Usage
evaluate.genome(PopName, allele, chromosome=NULL, begin=NULL, end=NULL,
verbose=FALSE)
Arguments
PopName |
Name of the population to be evaluated |
allele |
Function argument. |
chromosome |
Chromosome |
begin |
Start of chromosome region to investigate |
end |
End of the chromosome region |
verbose |
Function argument. |
Details
Analysis is done on the level of the chromosomes, independently of the loaded marker map.
Value
A data.frame with one row per individual. It always contains ratio, the evaluated allele proportion. With verbose = TRUE it additionally contains length, total, and noBlocks, as returned by the genome-evaluation routine.
Note
See the lib00.example1 for an application of the command.
evaluate.genotype()
Description
Calculates genotype frequencies
Usage
evaluate.genotype(PopName, Loci, alleles=0, mode=2, bg=0)
Arguments
PopName |
Name of the population to be evaluated |
Loci |
A List of Loci |
alleles |
A list of different Alleles |
mode |
Considering gametic phase or not |
bg |
Function argument. |
Value
A data.frame containing the genotype combinations and the count/frequency information returned by the compiled genotype evaluator. If a specific allele combination is requested the result can be reduced to the matching row. Returns NULL invisibly if no result is available.
Evaluate genotype information for selected loci in a population
Description
Evaluate genotype information for selected loci in a population.
Usage
evaluate.genotype2(PopName, Loci, alleles=0, mode=2, bg=0)
Arguments
PopName |
Name of the population. |
Loci |
Locus or loci to process. |
alleles |
Allele specification or coding. |
mode |
Evaluation or operation mode. |
bg |
Background or missing-data code used by the evaluator. |
Value
A data.frame containing genotype combinations and the count/frequency information returned by the compiled genotype evaluator. Returns NULL invisibly if no result is available.
Calculate haplotype information for a pair of loci
Description
Calculate haplotype information for a pair of loci.
Usage
evaluate.hap.calc(locusA, locusB)
Arguments
locusA |
First locus. |
locusB |
Second locus. |
Value
A numeric vector of length five, in the order noTI, noI, nodH, dLocA, and dLocB, containing the values returned by the compiled pairwise haplotype calculation.
Calculate a haplotype statistic for a pair of loci
Description
Calculate a haplotype statistic for a pair of loci.
Usage
evaluate.hap.calc.d(locusA, locusB)
Arguments
locusA |
First locus. |
locusB |
Second locus. |
Value
An invisible numeric vector of length three, in the order dd, dp, and rr, containing the values returned by the compiled haplotype-statistic calculation.
Release or reset temporary haplotype-evaluation state
Description
Release or reset temporary haplotype-evaluation state.
Usage
evaluate.hap.free()
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of releasing haplotype evaluation state.
Initialize haplotype evaluation for a population
Description
Initialize haplotype evaluation for a population.
Usage
evaluate.hap.init(PopName, doubleHetero=0, missing=0)
Arguments
PopName |
Name of the population. |
doubleHetero |
Control for handling double heterozygotes. |
missing |
Value or code used for missing data. |
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of initializing haplotype evaluation state.
Return genotype results from the current haplotype evaluation
Description
Return genotype results from the current haplotype evaluation.
Usage
evaluate.hap.return.genotype()
Value
A data.frame containing the genotype table from the current initialized haplotype evaluation.
Return the table produced by the current haplotype evaluation
Description
Return the table produced by the current haplotype evaluation.
Usage
evaluate.hap.return.table()
Value
A data.frame containing the table from the current initialized haplotype evaluation.
Evaluate haplotypes for selected loci in a population
Description
Evaluate haplotypes for selected loci in a population.
Usage
evaluate.haplotype(PopName, Loci, missing=0,
bg=plabsim$allele.missing.indicator)
Arguments
PopName |
Name of the population. |
Loci |
Locus or loci to process. |
missing |
Value or code used for missing data. |
bg |
Background or missing-data code used by the evaluator. |
Value
A data.frame containing the haplotype-evaluation table returned for the requested population and loci.
Evaluate linkage disequilibrium for a population and genomic position
Description
Evaluate linkage disequilibrium for a population and genomic position.
Usage
evaluate.ld(PopName, allele, chrom, pos)
Arguments
PopName |
Name of the population. |
allele |
Allele specification. |
chrom |
Chromosome identifier. |
pos |
Position on the chromosome. |
Value
A one-column data.frame named length, containing the linkage-disequilibrium segment lengths returned for the requested population, allele and position.
Evaluate selected loci in a population
Description
Evaluate selected loci in a population.
Usage
evaluate.locus(PopName, eLoci)
Arguments
PopName |
Name of the population. |
eLoci |
Locus or loci to evaluate. |
Value
A data.frame with one row for each requested locus and columns homozygote, heterozygote, and missing, containing the corresponding genotype counts.
Evaluate marker-assisted donor-parent information for populations
Description
Evaluate marker-assisted donor-parent information for populations.
Usage
evaluate.mdp(NamePg, NameP1, NameP2, effectfile=NULL)
Arguments
NamePg |
Name of the population to evaluate. |
NameP1 |
Name of the first parental population. |
NameP2 |
Name of the second parental population. |
effectfile |
Effect file or effect specification. |
Value
When a specific effectfile is supplied, the result is the invisible implementation-level list returned by the compiled MDP evaluator. When effectfile = NULL, the function evaluates all active effects for side effects and returns NULL.
Evaluate genetic values for simulation populations
Description
Evaluates genetic values for all active individuals in one or more simulation populations using loaded effect definitions.
Usage
evaluate.population(PopNames, effName=NULL)
Arguments
PopNames |
Population name or population specification accepted by the simulation routines. |
effName |
Optional character string naming the effect used as the selection index. |
Details
The population must have a complete genotype representation for all active
individuals. Evaluation calculates the effect-specific genetic values for all
loaded effects. Element 0 is first calculated as the weighted sum of these
values. If effName is supplied, element 0 is then replaced by the
genetic value of the named effect, preserving the existing selection
behavior.
Evaluation is population-wide. Allocation failure clears the genetic-value category rather than leaving only some individuals evaluated.
Any stored phenotypic values for a population being re-evaluated are discarded before the new genetic values are calculated.
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of calculating/storing evaluation values for the requested populations.
See Also
genotype.population, phenotype.population,
get.population.gvalue
gd.allele.frequencies()
Description
Calculates the allele frequencies of the populations in a data set consisting of codominant molecular marker data.
Usage
gd.allele.frequencies(dta)
Arguments
dta |
Name of the dataset contaning the observed marker genotypes. The format of the dataset is illustrated by the datasets gd.dummy.populations and gd.maize.populations. |
Value
A data.frame of allele frequencies, with marker/allele combinations in rows and populations in columns.
gd.allow.zero.frequencies()
Description
If for a marker no allele was observed, this is a missing value if codominant markers are used. However, for dominant markers where each band is regarded as a marker with only one allele, markers where no allele are observed are not missing data.
Usage
gd.allow.zero.frequencies(allow=0)
Arguments
allow |
Codominant marker data (SSR or RFLP) is analyzed like the datasets
|
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of setting whether zero allele frequencies are allowed in the genetic-diversity routines.
gd.correct.missing()
Description
The function gd.correct.missing corrects missing marker data points in a data set. Two rules are applied: (1) If at least one allele was observed at a marker, missing values for this marker at other alleles are changed to zeros. (2) If at a marker no allele was observed but for all alleles only zeros are in the data set, the zeros are changed to missing data.
Usage
gd.correct.missing(dta)
Arguments
dta |
Name of the dataset contaning the observed marker genotypes. The format of the dataset is illustrated by the datasets gd.dummy.populations and gd.maize.populations. |
Value
A data.frame containing the corrected data, with the same row and column structure as the supplied data object.
gd.data.parameters()
Description
Determines the structure of a data set containing molecular marker data.
Usage
gd.data.parameters(dta)
Arguments
dta |
Name of the dataset contaning the observed marker genotypes. The format of the dataset is illustrated by the datasets gd.dummy.populations and gd.maize.populations. |
Value
A named list describing the genetic-diversity data object, with components colhead, rowhead, no.mar, no.all, no.pop, no.ind, ind.list, mar.list, and pop.list. These contain parsed column and row headings, marker/allele/population counts, individual indices, marker names, and population names, respectively.
Gd distance similarity
Description
Calculate genetic distance or similarity measures, with optional bootstrap inputs.
Usage
gd.distance.similarity(dta, measure, par, boot=0, no.pop.u=0, pop.u=0,
no.ind.u=0, ind.u=0, no.mar.u=0, mar.u=0)
Arguments
dta |
Input data object. |
measure |
Distance or similarity measure. |
par |
Parameters used by the selected method. |
boot |
Bootstrap control. |
no.pop.u |
Number of populations used for bootstrap sampling. |
pop.u |
Population bootstrap indices. |
no.ind.u |
Numbers of individuals used for bootstrap sampling. |
ind.u |
Individual bootstrap indices. |
no.mar.u |
Number of markers used for bootstrap sampling. |
mar.u |
Marker bootstrap indices. |
Value
A data.frame with columns OTU1, OTU2, and Measure, containing pairwise population comparisons. When resampling-based uncertainty is requested an additional SDev column contains the corresponding standard deviations.
Molecular marker data for several dummy populations
Description
This dataset illustrates how missing values are handled in the calculation of allele freuquencies.
Usage
data(gd.dummy.populations)
gd.genetic.distance()
Description
Calculates pairwise genetic distances between all pairs of populations in the dataset.
Usage
gd.genetic.distance(dta, measure="euc", par="dist", boot=0, no.pop.u=0,
pop.u=0, no.ind.u=0, ind.u=0, no.mar.u=0, mar.u=0)
Arguments
dta |
Name of the dataset containing the observed marker genotypes. The format of the dataset is illustrated by the datasets gd.dummy.populations and gd.maize.populations. | ||||||||||
measure |
| ||||||||||
par |
Parameters to be calculated:
| ||||||||||
boot |
Number of bootstrap runs for estimating the standard deviation of the estimates. | ||||||||||
no.pop.u |
Function argument. | ||||||||||
pop.u |
Function argument. | ||||||||||
no.ind.u |
Function argument. | ||||||||||
ind.u |
Function argument. | ||||||||||
no.mar.u |
Function argument. | ||||||||||
mar.u |
Function argument. |
Value
A data.frame with columns OTU1, OTU2, and Measure, containing the requested pairwise genetic distances; an additional SDev column is included when resampling-based standard deviations are requested. Returns NULL invisibly on validation failure.
Note
1.) The Rogers distance and the Modified Rogers distance are
standardized measures in the interval [0,1]. The standardization is
achieved by the divisions by m and by 2 in the respective equations.
This requests codominant markers! These measures can't be used with
dominant markers. However the Euclidean distance is not standardized and
can also be used with dominant marker data (such as the AFLP data in the
dataset gd.maize.aflp).
2.) For homozygous genotypes Rogers' Distance (Rogers 1972) equals the Nei-Li Distance (Nei and Li 1979), which is the same as 1 minus the Dice coefficient (Dice 1945), further the Modified Rogers' Distance is the square root of the Rogers' Distance.
3.)
For calculating pairwise genetic distances between individuals you can
use populations consisting of one single individual each. The column
names could be e.g., p1.1 p2.1 p3.1 p4.1
References
Dice, L.R. 1945. Measures of the amount of ecologic association between species. Ecology 26:197–302.
Nei, M. and W.H. Li. 1979. Mathematical Models for studying genetic variation in terms of restriction endonucleases. Proc. Natl. Acad. Sci. USA. 76:5268–5371.
Rogers, J.S. 1972. Measures of genetic similarity and genetic distance. VII. Univ. Tex. Publ. 7213:145–153.
Wright, S. 1978. Evolution and the Genetics of Populations. Volume 4: Variability Within and Among Natural Populations. University of Chicago Press, Chicago, Illinois.
gd.list.irregular()
Description
The function gd.list.irregular lists individuals for which at one marker more than two alleles were observed.
Usage
gd.list.irregular(dta1)
Arguments
dta1 |
Name of the dataset containing the observed marker genotypes. The format of the dataset is illustrated by the datasets gd.dummy.populations and gd.maize.populations. |
Value
A data.frame with columns ind and marker, identifying individual/marker combinations for which more than two alleles were observed.
gd.list.missing()
Description
The function gd.list.missing lists the missing marker data points in a data set containing molecular marker data.
Usage
gd.list.missing(dta)
Arguments
dta |
Name of the dataset containing the observed marker genotypes. The format of the dataset is illustrated by the datasets gd.dummy.populations and gd.maize.populations. |
Value
A data.frame with columns ind and marker, identifying individual/marker combinations with missing observations.
51 maize lines analyzed with 462 AFLP markers.
Description
This data set consists of AFLP data from 51 maize lines. 462 markers
are used. This is dominant data. Each band is a marker, which has only
one allele. For analyzing this type of dominant marker use
gd.allow.zero.frequencies(1)
Usage
data(gd.maize.aflp)
Format
The first five rows and columns look like:
| l1.1 | l2.1 | l3.1 | l4.1 | l5.1 | |
| E35M50+M001.1 | 1 | 1 | 1 | 0 | 1 |
| E35M50+M002.1 | 0 | 0 | 0 | 0 | 1 |
| E35M50+M003.1 | 0 | 1 | 0 | 0 | 0 |
| E35M50+M004.1 | 1 | 1 | 1 | 0 | 1 |
| E35M50+M005.1 | 0 | 0 | 1 | 0 | 1 |
Each column contains the marker genotype of an individual. The identifier for an individual (the column name) contains the name of the population to which the individual belongs and the name of the individual, separated by a dot. For band there is a row in the dataset, the rowname contains the name of the marker and seperated by a dot a 1, indicating that each marker has only one allele.
The data matrix consists of zeros and ones, denoting whether a certain band was observed at an individual or not. Missing data are denoted with NA.
For AFLPs the primer combination and the band is separated by a +.
See Also
gd.genetic.distance,
gd.similarity.coefficient,
gd.maize.populations,
gd.maize.lines,
gd.salad.aflp
RFLP data from 50 maize lines
Description
This data set consists of the RFLP data from 50 maize lines.
Usage
data(gd.maize.lines)
Format
The format is described for the data set gd.maize.populations.
See Also
gd.genetic.distance,
gd.similarity.coefficient,
gd.maize.populations,
gd.maize.aflp,
gd.salad.aflp
SSR data from seven maize populations
Description
This data set consists of the SSR marker data of seven tropical maize populations. From each populations 48 individuals were samples analyzed with 85 polymorphic Single Sequence Repeat markers.
Usage
data(gd.maize.populations)
Format
The first five rows and columns look like:
| POL24.01 | POL24.02 | POL24.03 | POL24.04 | POL25.01 | |
| nc130.139 | 0 | 0 | 1 | 1 | NA |
| nc130.142 | 1 | 1 | 0 | 1 | 1 |
| nc130.145 | 0 | 1 | 0 | 0 | 0 |
| nc133.110 | 1 | 0 | 0 | 1 | 1 |
| nc133.148 | 0 | 0 | 0 | 0 | 0 |
Each column contains the marker genotype of an individual. The identifier for an individual (the column name) contains the name of the population to which the individual belongs and the name of the individual, separated by a dot. For each allel of a marker there is a row in the dataset, the rowname contains the name of the marker and the name of the allele seperated by a dot.
The data matrix consists of zeros and ones, denoting whether a certain allale was observed at an individual or not. Missing data are denoted with NA.
See Also
gd.genetic.distance,
gd.similarity.coefficient,
gd.maize.lines,
gd.maize.aflp,
gd.salad.aflp
Construct a matrix from genetic-distance results
Description
Construct a matrix from genetic-distance results.
Usage
gd.mk.matr(dta, distance)
Arguments
dta |
Input data object. |
distance |
Distance specification or matrix. |
Value
A numeric matrix representation of the supplied pairwise distance results, with population labels taken from the data object.
gd.pcoa()
Description
Calculates a principal coordinate analysis for marker
data. It is a wrapper for the cmdscale function.
Usage
gd.pcoa(dta, k=3)
Arguments
dta |
Name of the dataset containing the observed marker genotypes |
k |
Number of the principal coordinates that should be returned |
Value
A numeric matrix of principal-coordinate scores returned by cmdscale(), with rows corresponding to populations and columns to the first k principal coordinates.
Jasminas 44 salads analyzed with 108 AFLP markers.
Description
This data set consists of AFLP data from 44 salad lines. 108 markers
are used. This is dominant data. Each band is a marker, which has only
one allele. For analyzing this type of dominant marker use
gd.allow.zero.frequencies(1)
Usage
data(gd.salad.aflp)
Format
The first five rows and columns look like:
| l34.1 | l44.1 | l45.1 | l38.1 | l33.1 | |
| p1+6.1 | 1 | 0 | 0 | 1 | 1 |
| p2+1.1 | 1 | 1 | 1 | 1 | 1 |
| p2+6.1 | 1 | 1 | 1 | 1 | 1 |
| p2+7.1 | 1 | 1 | 1 | 1 | 1 |
| p2+8.1 | 1 | 1 | 1 | 1 | 1 |
Each column contains the marker genotype of an individual. The identifier for an individual (the column name) contains the name of the population to which the individual belongs and the name of the individual, separated by a dot. For band there is a row in the dataset, the rowname contains the name of the marker and seperated by a dot a 1, indicating that each marker has only one allele.
The data matrix consists of zeroes and ones, denoting whether a certain band was observed at an individual or not. Missing data are denoted with NA.
For AFLPs the primer combination and the band is separated by a +.
See Also
gd.genetic.distance,
gd.similarity.coefficient,
gd.maize.populations,
gd.maize.lines,
gd.maize.aflp
gd.similarity.coefficient()
Description
Calculates matching coefficients between individuals.
Usage
gd.similarity.coefficient(dta, measure="jac", par="dist", boot=0, no.pop.u=0,
pop.u=0, no.ind.u=0, ind.u=0, no.mar.u=0, mar.u=0,
resample.primers=FALSE)
Arguments
dta |
Name of the dataset containing the observed marker genotypes. The format of the dataset is illustrated by the dataset gd.salad.aflp. | ||||||
measure |
| ||||||
par |
Parameters to be calculated:
| ||||||
boot |
Number of bootstrap runs for estimating the standard deviation of the estimates. | ||||||
no.pop.u |
Function argument. | ||||||
pop.u |
Function argument. | ||||||
no.ind.u |
Function argument. | ||||||
ind.u |
Function argument. | ||||||
no.mar.u |
Function argument. | ||||||
mar.u |
Function argument. | ||||||
resample.primers |
if set to |
Value
A data.frame with columns OTU1, OTU2, and Measure, containing the requested pairwise similarity coefficients; an additional SDev column is included when resampling-based standard deviations are requested. Returns NULL invisibly on validation failure.
Note
Each population should consist of one indivdiual only.
Split or transform marker labels used by genetic-diversity routines
Description
Split or transform marker labels used by genetic-diversity routines.
Usage
gd.splitdot(m.a)
Arguments
m.a |
Marker or matrix input. |
Value
A character matrix with two columns and one row for each input string. The first column contains the part before the first dot (or underscore if no dot is present), and the second column contains the remaining part; if no separator is present the second field is empty.
Return status information concerning zero allele frequencies
Description
Return status information concerning zero allele frequencies.
Usage
gd.status.zero.frequencies()
Value
An integer scalar giving the current compiled status value for whether zero allele frequencies are allowed.
generate.effect.file()
Description
Generates a file with effect descriptions that can be used for selection. Effects can be loaded. Additive and dominance effects can be
Usage
generate.effect.file(fName, description)
Arguments
fName |
Name of the file |
description |
Description of the effect |
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of generating the effect file.
Note
See define.effects for the syntax of the
descripton and for an example
generate.map.file()
Description
Generates a simulation linkage-map text file from a compact map description without itself installing the map.
Usage
generate.map.file(fName, description)
Arguments
fName |
The name of the map file that is generated |
description |
Description of the linkage map |
Details
The description syntax is the same syntax consumed by define.map. The function validates the file name and description, normalizes a multi-element description to one comma-separated specification, and asks the simulation backend to write the map file.
The generated file can subsequently be loaded as a simulation linkage map. Generating the file does not by itself replace the currently loaded map.
Value
An invisible list whose retval component is the compiled map-generation status. Invalid map dimensions return list(retval = -2L) invisibly. The main effect is writing the generated map file.
generate.population()
Description
Generates one or more simulation populations from the marker-incidence matrix representation used by the genetic-distance routines.
Usage
generate.population(dta, backcross=FALSE)
Arguments
dta |
Data frame or matrix in the marker-incidence representation expected by the genetic-distance population interface. |
backcross |
Logical flag selecting backcross-oriented generation in the backend. |
Details
The function interprets the row and column descriptors of dta to determine marker names, allele codes, population names, and individual membership. Allele coding must be numeric integer coding compatible with the simulation routines.
The R layer derives the population and marker dimensions and passes the matrix together with allele and missing-value information to the C population generator. The resulting objects are simulation populations represented by chromosome data, not SelectionTools marker-data sets.
The backcross flag selects the corresponding population-generation mode in the simulation backend. Existing genome and linkage-map definitions determine how generated marker information is represented in simulated chromosomes.
Value
An invisible implementation-level list returned by either the population-generation routine or the information routine used to report invalid allele encoding. The meaningful result is the side effect of generating the requested population when the input is valid.
genome.contribution()
Description
Calculates the distribution of one allele in populations.
Usage
genome.contribution(pops, allele, chromosome=NULL, begin=NULL, end=NULL)
Arguments
pops |
Names of the populations separated by blanks |
allele |
Allele which is considered |
chromosome |
Chromosome |
begin |
Start of chromosome region to investigate |
end |
End of the chromosome region |
Details
Analysis is done on the level of the chromosomes, independently of the
loaded marker map. Summarizes the results of evaluate.genome.
Value
A numeric matrix with one row per population and columns Obs, Mean, SDev, Min, Q10, Med, and Max, summarizing genome-contribution values.
Note
See the lib00.example1 for an application of the command.
Return the current genome parameters
Description
Returns the chromosome structure currently installed in the simulation backend.
Usage
genome.parameter.get()
Details
The result describes the simulation-wide genome parameters used to interpret chromosome segments and linkage-map positions. It includes the number of chromosomes, the number of homologues, and the chromosome-length vector.
Value
A named list with components no.chrom (number of chromosomes), no.hom (number of homologues), and chrom.len (chromosome-length vector).
Set genome parameters used by the simulation code
Description
Defines the chromosome structure used by the simulation backend.
Usage
genome.parameter.set(no.chrom, no.hom=2, chrom.len)
Arguments
no.chrom |
Positive number of chromosomes. |
no.hom |
Number of homologues; the default is two. |
chrom.len |
Numeric vector containing one chromosome length for each chromosome. |
Details
The genome definition is simulation-wide and determines the chromosome and homologue structure expected by population initialization, meiosis, crossing, single-seed descent, doubled-haploid production, and linkage-map operations.
chrom.len supplies one length per chromosome. Operations that model ordinary diploid marker data use two homologues. Changing genome parameters should therefore be coordinated with the loaded simulation linkage map and existing simulation populations.
Value
An invisible implementation-level list returned by the compiled state-setting or information routine. The function is primarily called for its side effect of resetting populations/map as needed and setting the genome parameters.
Extract or summarize genome segments for populations
Description
Extract or summarize genome segments for populations.
Usage
genome.segments(pops, allele, chromosome=NULL, begin=NULL, end=NULL)
Arguments
pops |
Population name or names. |
allele |
Allele specification. |
chromosome |
Chromosome identifier. |
begin |
Beginning position of a region. |
end |
Ending position of a region. |
Value
A numeric matrix with one row per population and columns Obs, Mean, SDev, Min, Q10, Med, and Max, summarizing the number of genome segments.
genotype.population()
Description
Calculates marker genotypes for one or more simulation populations from their chromosome-segment representation and the currently loaded linkage map.
Usage
genotype.population(PopNames)
Arguments
PopNames |
One population name, a space-separated character string of population names, or a character vector of population names to genotype. |
Details
Simulation populations store chromosome segments as the primary genetic representation. This function projects those segments onto the loci of the currently loaded map and stores the two alleles for every active individual and mapped locus.
All active individuals of a population are treated as one genotype category: after a successful call every active individual has genotype data. If a population contains a mixture of genotyped and non-genotyped active individuals, the complete population is recalculated. Allocation failure clears the population genotype cache rather than leaving a partial mixture.
Population names may be supplied as a vector; the R wrapper combines multiple names into the space-separated form accepted by the C routine. Genotyping does not evaluate genetic values. Existing evaluations that depend on a genotype may need to be recalculated after changes to the underlying population or map.
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of generating/storing genotype information for the requested populations.
get.genome.par()
Description
Returns the current simulation genome parameters.
Usage
get.genome.par()
Details
The returned values describe the genome used by simulation populations: chromosome count, number of homologues, and chromosome lengths. These parameters are separate from a SelectionTools marker data set, although the marker-to-simulation bridge can derive and install compatible simulation genome parameters from a marker map.
Value
Returns the same named list as genome.parameter.get(), with components no.chrom, no.hom, and chrom.len.
get.map()
Description
Returns the currently loaded simulation linkage map.
Usage
get.map()
Details
The result describes the simulation map points in their active access order, including chromosome, map position, locus name, and class. Retrieving the map does not modify simulation populations or marker-data sets.
Value
Returns the same data.frame as linkage.map.get(), with columns chrom, pos, name, and class.
Return the current marker-assisted donor-parent information
Description
Return the current marker-assisted donor-parent information.
Usage
get.mdp()
Value
An integer scalar containing the current marker-assisted donor-parent (MDP) value returned by the compiled routine.
get.population()
Description
Returns the raw chromosome-segment representation of a simulation population as an R data frame.
Usage
get.population(PopName)
Arguments
PopName |
Character string naming the simulation population. |
Details
Each row identifies an individual, chromosome, homologue, segment start position, and allele. These rows are the primary simulation representation from which marker genotypes are derived when genotype.population is run.
The returned data frame can be used to inspect or serialize a population and is the representation accepted by init.population. Retrieving the population does not alter genotype or evaluation caches.
Value
A data.frame with columns ind, chrom, hom, pos, and all, representing the active population in the package population format.
Get genetic values from a simulation population
Description
Returns stored evaluated genetic values for the active individuals of a simulation population.
Usage
get.population.gvalue(name, EffName=NULL)
Arguments
name |
Character string naming the population. |
EffName |
Optional character string identifying the effect-specific values to retrieve. |
Details
Genetic values are derived population data created by population evaluation
or explicitly supplied through set.population.gvalue.
If EffName is omitted, element 0 of the genetic value array is
returned. After ordinary population evaluation this is the weighted genetic
selection index. If EffName is supplied, the corresponding
effect-specific genetic value is returned.
Value
A one-column data.frame named gvalue, containing one genetic value per active individual in the requested population.
See Also
evaluate.population, set.population.gvalue,
get.population.pvalue
Return stored information for individuals in a population
Description
Return stored information for individuals in a population.
Usage
get.population.info(PopName, ind)
Arguments
PopName |
Name of the population. |
ind |
Individual index or indices. |
Value
Returns the same character value or NULL as population.info.get().
Get phenotypic values from a simulation population
Description
Returns stored phenotypic values for the active individuals of a simulation population.
Usage
get.population.pvalue(name, EffName=NULL)
Arguments
name |
Character string naming the population. |
EffName |
Optional character string naming an effect. |
Details
Phenotypic values must previously have been generated with
phenotype.population. If EffName is omitted, element 0 of the
phenotypic value array is returned. This is the weighted phenotypic selection
index. If EffName is supplied, the corresponding effect-specific
phenotypic value is returned.
Value
A one-column data.frame named pvalue, containing one phenotypic value per active individual in the requested population.
See Also
phenotype.population, get.population.gvalue
get.population.size()
Description
Returns the active size and allocated capacity of a simulation population.
Usage
get.population.size(PopName)
Arguments
PopName |
Name of the population |
Details
The active size is the number of individuals currently belonging to the population. The allocated capacity is internal storage available to the population and can differ from the active size after resizing or other population operations.
Value
Returns the same one-row data.frame as population.size.get(), with columns NoInds and NoIndsAlloc, or NULL if the population is unavailable.
get.score()
Description
The function get.score() returns a field with the calculated genotypic index for every individual of the population pop. If an effect is specified then the value of this effect is returned.
Usage
get.score(pop, effectfile=NULL)
Arguments
pop |
Name of the evaluated population. |
effectfile |
Name of effect which should be evaluated. |
Value
A one-column data.frame named gvalue, containing the genetic values obtained after genotyping and evaluating the requested population.
Build the genomic-selection variance or covariance matrix for a data set
Description
Build the genomic-selection variance or covariance matrix for a data set.
Usage
gs.build.V(out.filename = "V.matrix", auxfiles = FALSE, data.set = "default")
Arguments
out.filename |
Character string retained for compatibility and used as part of the name of an internal file in the R session temporary directory. The temporary file is removed before the function returns. |
auxfiles |
Logical. If |
data.set |
Name of the SelectionTools data set. |
Value
Invisibly returns a numeric matrix containing the generated variance/covariance matrix V when auxfiles = TRUE and the compiled calculation succeeds; otherwise NULL.
gs.build.Z()
Description
The function gs.build.Z builds the design (Z) matrix out of the marker data for genomic selection.
Usage
gs.build.Z(out.filename = "Z.matrix", auxfiles = FALSE, data.set = "default")
Arguments
out.filename |
Character string retained for compatibility and used as part of the name of an internal file in the R session temporary directory. The temporary file is removed before the function returns. |
auxfiles |
Logical. If |
data.set |
character. Name of the data set that is used by the function |
Value
Invisibly returns a matrix containing the generated marker/design matrix Z when auxfiles = TRUE and generation succeeds; otherwise NULL.
Build estimation and validation subsets from a base data set
Description
Build estimation and validation subsets from a base data set.
Usage
gs.build.esvs(es.size, base.set = "default", estimation.set = "none",
validation.set = "none", es_m = NULL, es_p = NULL,
vs_m = NULL, vs_p = NULL, auxfiles = FALSE)
Arguments
es.size |
Size of the estimation set. |
base.set |
Name of the base data set. |
estimation.set |
Name of the estimation data set. |
validation.set |
Name of the validation data set. |
es_m |
Marker-data output file for the estimation set. Required when |
es_p |
Phenotype-data output file for the estimation set. Required when |
vs_m |
Marker-data output file for the validation set. Required when |
vs_p |
Phenotype-data output file for the validation set. Required when |
auxfiles |
Logical. If |
Value
On a successful run, an invisible implementation-level list returned by the final timing/information helper; the meaningful result is the side effect of constructing the estimation and validation subsets. Validation failures may return NULL.
Build training and prediction subsets from a base data set
Description
Build training and prediction subsets from a base data set.
Usage
gs.build.tsps(ts.size, base.set = "default", training.set = "none",
prediction.set = "none", es_m = NULL, es_p = NULL,
vs_m = NULL, vs_p = NULL, auxfiles = FALSE)
Arguments
ts.size |
Size of the training set. |
base.set |
Name of the base data set. |
training.set |
Name of the training data set. |
prediction.set |
Name of the prediction data set. |
es_m |
Marker-data output file for the training set. Required when |
es_p |
Phenotype-data output file for the training set. Required when |
vs_m |
Marker-data output file for the prediction set. Required when |
vs_p |
Phenotype-data output file for the prediction set. Required when |
auxfiles |
Logical. If |
Value
On a successful run, an invisible implementation-level list returned by the final timing/information helper; the meaningful result is the side effect of constructing the training and prediction subsets. Validation failures may return NULL.
Check whether p-values are available for a genomic-selection data set
Description
Check whether p-values are available for a genomic-selection data set.
Usage
gs.check.pvals(data.set ="default")
Arguments
data.set |
Name of the SelectionTools data set. |
Value
A logical scalar indicating whether p-values are available for the requested genomic-selection data set.
gs.compare.effects()
Description
Creates a scatter plot of effects of two different models to visualize differences in estimated effects.
Usage
gs.compare.effects(data.set.a, data.set.b, label.a = data.set.a,
label.b = data.set.b)
Arguments
data.set.a |
character. Name of the first data set containing marker effects. These effects are plotted on the x axis. |
data.set.b |
character. Name of the second data set containing marker effects. These effects are plotted on the y axis. |
label.a |
character. Label of the x axis. |
label.b |
character. Label of the y axis. |
Value
No useful return value; the function is called for its side effect of comparing and plotting marker-effect results. Validation failures return NULL invisibly.
Evaluate crosses using expected selection response
Description
Evaluate crosses using expected selection response.
Usage
gs.cross.eval.es(data.set = "default", alpha = 0.1, N = 0, G = 0)
Arguments
data.set |
Name of the SelectionTools data set. |
alpha |
Method-specific tuning or significance parameter. |
N |
Population-size or method-specific size parameter. |
G |
Generation or method-specific generation parameter. |
Value
On a successful run, an invisible implementation-level list returned by the final timing/information helper; the meaningful result is the side effect of evaluating crosses by expected selection response. Validation failures may return NULL.
Evaluate crosses using genetic distance
Description
Evaluate crosses using genetic distance.
Usage
gs.cross.eval.gd(dist = "rd", data.set = "default")
Arguments
dist |
Genetic-distance measure. |
data.set |
Name of the SelectionTools data set. |
Value
On a successful run, an invisible implementation-level list returned by the final timing/information helper; the meaningful result is the side effect of evaluating crosses by genetic distance. Validation failures may return NULL.
Evaluate crosses using genetic-distance function
Description
Evaluate crosses using genetic-distance function.
Usage
gs.cross.eval.gd.fct(dist = "rd", split = 1, data.set = "default")
Arguments
dist |
Genetic-distance measure. |
split |
Split or partition control. |
data.set |
Name of the SelectionTools data set. |
Value
On a successful run, an invisible implementation-level list returned by the final timing/information helper; the meaningful result is the side effect of evaluating factorial crosses by genetic distance. Validation failures may return NULL.
Evaluate crosses using marker effects
Description
Evaluate crosses using marker effects.
Usage
gs.cross.eval.ma(data.set = "default")
Arguments
data.set |
Name of the SelectionTools data set. |
Value
On a successful run, an invisible implementation-level list returned by the final timing/information helper; the meaningful result is the side effect of evaluating crosses using marker effects. Validation failures may return NULL.
Evaluate crosses using marker information
Description
Evaluate crosses using marker information.
Usage
gs.cross.eval.mi(data.set = "default")
Arguments
data.set |
Name of the SelectionTools data set. |
Value
On a successful run, an invisible implementation-level list returned by the final timing/information helper; the meaningful result is the side effect of evaluating crosses by marker information. Validation failures may return NULL.
Evaluate crosses using cross mean
Description
Evaluate crosses using cross mean.
Usage
gs.cross.eval.mu(data.set = "default")
Arguments
data.set |
Name of the SelectionTools data set. |
Value
On a successful run, an invisible implementation-level list returned by the final timing/information helper; the meaningful result is the side effect of evaluating crosses by expected cross mean. Validation failures may return NULL.
Evaluate crosses using cross variance
Description
Evaluate crosses using cross variance.
Usage
gs.cross.eval.va(data.set = "default", pop.type = "unlinked", t = 0,
map.function = "Haldane")
Arguments
data.set |
Name of the SelectionTools data set. |
pop.type |
Population type. |
t |
Generation or method-specific time parameter. |
map.function |
Map function to use. |
Value
On a successful run, an invisible implementation-level list returned by the final timing/information helper; the meaningful result is the side effect of evaluating crosses by cross variance. Validation failures may return NULL.
Return or write information for evaluated crosses
Description
Return or write information for evaluated crosses.
Usage
gs.cross.info(bestn = 0, sortby = "index", out.filename ="cross.info",
data.set ="default")
Arguments
bestn |
Number of best entries to retain. |
sortby |
Criterion used for sorting. |
out.filename |
Character string retained for compatibility and used as part of the name of an internal file in the R session temporary directory. The temporary file is removed before the function returns. |
data.set |
Name of the SelectionTools data set. |
Value
Invisibly returns a data.frame containing the evaluated-cross information read from the generated result when the compiled operation succeeds; otherwise NULL.
Return or write genetic-distance information for crosses
Description
Return or write genetic-distance information for crosses.
Usage
gs.cross.info.gd(out.filename ="cross.info.gd", data.set ="default")
Arguments
out.filename |
Character string retained for compatibility and used as part of the name of an internal file in the R session temporary directory. The temporary file is removed before the function returns. |
data.set |
Name of the SelectionTools data set. |
Value
Invisibly returns a data.frame containing the genetic-distance cross information read from the generated result when the compiled operation succeeds; otherwise NULL.
gs.cross.validation()
Description
Splits a given data set into estimation and prediction set. Estimates marker effects with the specified model using the estimation set and predicts direct genomic values (DGV) for the built prediction set.
Usage
gs.cross.validation(estimation.method, n.ts, n.runs, hsq = 0.8, alpha = 1,
maxiter = 100, precision = 0.0001,
out.filename.r = "crosval", out.filename.u = "effects",
auxfiles = TRUE, data.set = "default")
Arguments
estimation.method |
character. Specifies the model: "rmlc", "rrwe", "rmlv", "rmlr", "rrwr" "bayes" |
n.ts |
Number of individuals in the training set |
n.runs |
Number of cross validation runs |
hsq |
Heritability of the analyzed trait. Used for "rrwe", "rrwr" |
alpha |
Function argument. |
maxiter |
Maximum number of iterations used for convergence of the REML estimator. Used if scheme "rmlc" or "rmlv" |
precision |
Gives the precision as second stop value for iteration next to maxiter. If the difference of the values of lambda in two subsequent iterations is less than the given precision, convergence is reached and the iteration will stop. Used if scheme "rmlc" or "rmlv" |
out.filename.r |
Character string retained for compatibility and used as part of the name of the internal session-temporary file used for cross-validation correlations. |
out.filename.u |
Character string retained for compatibility and used as part of the name of the internal session-temporary file used for estimated marker effects. |
auxfiles |
Logical. If |
data.set |
character. Name of the data set that is used by the function. |
Value
Invisibly returns a list when auxfiles = TRUE and cross-validation succeeds. Component cor is a numeric matrix of cross-validation correlations and component eff is a numeric matrix of estimated effects. Otherwise NULL.
Estimate marker effects using the LSQ method
Description
Estimate marker effects using the LSQ method.
Usage
gs.esteff.lsq(out.filename = "eff.lsq", auxfiles = TRUE, data.set ="default")
Arguments
out.filename |
Character string retained for compatibility and used as part of the name of an internal file in the R session temporary directory. The temporary file is removed before the function returns. |
auxfiles |
Logical. If |
data.set |
Name of the SelectionTools data set. |
Value
Invisibly returns a numeric matrix containing the least-squares marker-effect estimates when auxfiles = TRUE and the compiled calculation succeeds; otherwise NULL.
Estimate marker effects using the RMLA method
Description
Estimate marker effects using the RMLA method.
Usage
gs.esteff.rmla(alpha = 1, maxiter = 1000, precision = 0.0001, hsq = 0.9,
out.filename = "eff.rmla", auxfiles = TRUE,
data.set ="default")
Arguments
alpha |
Method-specific tuning or significance parameter. |
maxiter |
Maximum number of iterations. |
precision |
Convergence tolerance. |
hsq |
Heritability parameter. |
out.filename |
Character string retained for compatibility and used as part of the name of an internal file in the R session temporary directory. The temporary file is removed before the function returns. |
auxfiles |
Logical. If |
data.set |
Name of the SelectionTools data set. |
Value
Invisibly returns a numeric matrix containing the RMLA marker-effect estimates when auxfiles = TRUE and the compiled calculation succeeds; otherwise NULL.
gs.esteff.rmlc()
Description
Estimates marker effects using BLUP with constant variances for each marker.
Usage
gs.esteff.rmlc(maxiter = 1000, precision = 0.0001, hsq = 0.9,
out.filename = "eff.rmlc", auxfiles = FALSE,
data.set ="default")
Arguments
maxiter |
Maximum number of iterations used for convergence of the REML estimator |
precision |
Gives the precision as second stop value for iteration next to maxiter. If the difference of the values of lambda in two subsequent iterations is less than the given precision, convergence is reached and the iteration will stop |
hsq |
Function argument. |
out.filename |
Character string retained for compatibility and used as part of the name of an internal file in the R session temporary directory. The temporary file is removed before the function returns. |
auxfiles |
Logical. If |
data.set |
character. Name of the data set that is used by the function |
Details
The restricted maximum likelihood (REML) method is used to estimate variance components. In a next step, lambda is calculated using the formula:
\lambda=\frac{\sigma_e^2}{\frac{\sigma_g^2}{m}}
where \sigma_e^2 is the residual error variance, \sigma_g^2 is the
genetic variance and m is the number of markers.
Value
Invisibly returns a numeric matrix containing the RMLC marker-effect estimates when auxfiles = TRUE and the compiled calculation succeeds; otherwise NULL.
Estimate marker effects using the RMLR method
Description
Estimate marker effects using the RMLR method.
Usage
gs.esteff.rmlr(maxiter = 1000, precision = 0.001, hsq = 0.9,
out.filename = "eff.rmlr", auxfiles = TRUE,
data.set ="default")
Arguments
maxiter |
Maximum number of iterations. |
precision |
Convergence tolerance. |
hsq |
Heritability parameter. |
out.filename |
Character string retained for compatibility and used as part of the name of an internal file in the R session temporary directory. The temporary file is removed before the function returns. |
auxfiles |
Logical. If |
data.set |
Name of the SelectionTools data set. |
Value
Invisibly returns a numeric matrix containing the RMLR marker-effect estimates when auxfiles = TRUE and the compiled calculation succeeds; otherwise NULL.
gs.esteff.rmlv()
Description
Estimates marker effects using BLUP with variable variances for the markers.
Usage
gs.esteff.rmlv(maxiter = 1000, precision = 0.001, hsq = 0.9,
out.filename = "eff.rmlv", auxfiles = FALSE,
data.set ="default")
Arguments
maxiter |
Maximum number of iterations used for convergence of the REML estimator. |
precision |
Gives the precision as second stop value for iteration next to maxiter. If the difference of the values of lambda in two subsequent iterations is less than the given precision, convergence is reached and the iteration will stop. |
hsq |
Function argument. |
out.filename |
Character string retained for compatibility and used as part of the name of an internal file in the R session temporary directory. The temporary file is removed before the function returns. |
auxfiles |
Logical. If |
data.set |
character. Name of the data set that is used by the function. |
Details
The restricted maximum likelihood (REML) method is used to estimate variance components.
Value
Invisibly returns a numeric matrix containing the RMLV marker-effect estimates when auxfiles = TRUE and the compiled calculation succeeds; otherwise NULL.
Estimate marker effects using the RR method
Description
Estimate marker effects using the RR method.
Usage
gs.esteff.rr(method = "BLUP", maxiter = 1000, precision = 0.0001, hsq = 0.80,
alpha = 1, out.filename = "eff.rr", auxfiles = FALSE,
data.set = "default")
Arguments
method |
Method to use. |
maxiter |
Maximum number of iterations. |
precision |
Convergence tolerance. |
hsq |
Heritability parameter. |
alpha |
Method-specific tuning or significance parameter. |
out.filename |
Character string retained for compatibility and used as part of the name of an internal file in the R session temporary directory. The temporary file is removed before the function returns. |
auxfiles |
Logical. If |
data.set |
Name of the SelectionTools data set. |
Value
Invisibly returns a numeric matrix containing the estimated marker effects from ridge regression when auxfiles = TRUE and the compiled calculation succeeds; otherwise NULL.
Estimate genomic values for an estimation or prediction data set
Description
Estimate genomic values for an estimation or prediction data set.
Usage
gs.estimate.gv(estimation.set ="default", prediction.set ="default",
out.filename = "breeding.values", auxfiles = TRUE)
Arguments
estimation.set |
Name of the estimation data set. |
prediction.set |
Name of the prediction data set. |
out.filename |
Character string retained for compatibility and used as part of the name of an internal file in the R session temporary directory. The temporary file is removed before the function returns. |
auxfiles |
Logical. If |
Value
Invisibly returns a data.frame with genomic-value results when auxfiles = TRUE and estimation succeeds. It contains individual identifiers and predicted genomic values, and also observed phenotypes when those are present in the generated result. Otherwise NULL.
Return the genomic-selection variance or covariance matrix for a data set
Description
Return the genomic-selection variance or covariance matrix for a data set.
Usage
gs.get.V(data.set = "default")
Arguments
data.set |
Name of the SelectionTools data set. |
Value
Invisibly returns the current numeric variance/covariance matrix V; returns NULL if it cannot be obtained from the current data set.
Return the genomic-selection marker or design matrix for a data set
Description
Return the genomic-selection marker or design matrix for a data set.
Usage
gs.get.Z(data.set = "default")
Arguments
data.set |
Name of the SelectionTools data set. |
Value
Invisibly returns the current numeric marker/design matrix Z; returns NULL if it cannot be obtained from the current data set.
Return the genomic-selection information matrix for a data set
Description
Return the genomic-selection information matrix for a data set.
Usage
gs.get.im(data.set = "default")
Arguments
data.set |
Name of the SelectionTools data set. |
Value
An invisible named integer vector with components i, the number of individuals, and m, the number of markers in the requested data set.
Return the genomic-selection information level for a data set
Description
Return the genomic-selection information level for a data set.
Usage
gs.get.info.level(data.set ="default")
Arguments
data.set |
Name of the SelectionTools data set. |
Value
An integer scalar giving the current genomic-selection information level for the requested data set.
Return the genomic-selection response vector for a data set
Description
Return the genomic-selection response vector for a data set.
Usage
gs.get.y(data.set = "default")
Arguments
data.set |
Name of the SelectionTools data set. |
Value
Invisibly returns the current numeric response vector y; returns NULL if it cannot be obtained from the current data set.
Print or handle a genomic-selection information message
Description
Print or handle a genomic-selection information message.
Usage
gs.info(lev, msg)
Arguments
lev |
Information or verbosity level. |
msg |
Message text. |
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of issuing the requested genomic-selection information message.
Estimate or set genomic-selection lambda values using the aov method
Description
Estimate or set genomic-selection lambda values using the aov method.
Usage
gs.lambda.aov(hsq, alpha = 1, out.filename = "lambda", auxfiles = TRUE,
data.set = "default")
Arguments
hsq |
Heritability parameter. |
alpha |
Method-specific tuning or significance parameter. |
out.filename |
Character string retained for compatibility and used as part of the name of an internal file in the R session temporary directory. The temporary file is removed before the function returns. |
auxfiles |
Logical. If |
data.set |
Name of the SelectionTools data set. |
Value
Invisibly returns a numeric matrix containing the lambda results from the ANOVA method when auxfiles = TRUE and the compiled calculation succeeds; otherwise NULL.
Estimate or set genomic-selection lambda values using the const method
Description
Estimate or set genomic-selection lambda values using the const method.
Usage
gs.lambda.const(lambda, out.filename = "lambda", auxfiles = FALSE,
data.set = "default")
Arguments
lambda |
Regularization parameter. |
out.filename |
Character string retained for compatibility and used as part of the name of an internal file in the R session temporary directory. The temporary file is removed before the function returns. |
auxfiles |
Logical. If |
data.set |
Name of the SelectionTools data set. |
Value
Invisibly returns a numeric matrix containing the lambda results from the constant-lambda calculation when auxfiles = TRUE and the compiled calculation succeeds; otherwise NULL.
Estimate or set genomic-selection lambda values using the emstep method
Description
Estimate or set genomic-selection lambda values using the emstep method.
Usage
gs.lambda.emstep(constvar = FALSE, out.filename = "lambda", auxfiles = TRUE,
data.set = "default")
Arguments
constvar |
Logical flag controlling use of a constant variance. |
out.filename |
Character string retained for compatibility and used as part of the name of an internal file in the R session temporary directory. The temporary file is removed before the function returns. |
auxfiles |
Logical. If |
data.set |
Name of the SelectionTools data set. |
Value
Invisibly returns a numeric matrix containing the lambda results from the EM-step method when auxfiles = TRUE and the compiled calculation succeeds; otherwise NULL.
Estimate or set genomic-selection lambda values using the hsq method
Description
Estimate or set genomic-selection lambda values using the hsq method.
Usage
gs.lambda.hsq(hsq, out.filename = "lambda", auxfiles = TRUE,
data.set = "default")
Arguments
hsq |
Heritability parameter. |
out.filename |
Character string retained for compatibility and used as part of the name of an internal file in the R session temporary directory. The temporary file is removed before the function returns. |
auxfiles |
Logical. If |
data.set |
Name of the SelectionTools data set. |
Value
Invisibly returns a numeric matrix containing the lambda results calculated from heritability when auxfiles = TRUE and the compiled calculation succeeds; otherwise NULL.
Estimate or set genomic-selection lambda values using the reg method
Description
Estimate or set genomic-selection lambda values using the reg method.
Usage
gs.lambda.reg(hsq, out.filename = "lambda", auxfiles = TRUE,
data.set = "default")
Arguments
hsq |
Heritability parameter. |
out.filename |
Character string retained for compatibility and used as part of the name of an internal file in the R session temporary directory. The temporary file is removed before the function returns. |
auxfiles |
Logical. If |
data.set |
Name of the SelectionTools data set. |
Value
Invisibly returns a numeric matrix containing the lambda results from the regression method when auxfiles = TRUE and the compiled calculation succeeds; otherwise NULL.
Estimate or set genomic-selection lambda values using the rmla method
Description
Estimate or set genomic-selection lambda values using the rmla method.
Usage
gs.lambda.rmla(alpha = 1, maxiter = 1000, precision = 0.0001, hsq = 0.9,
out.filename = "lambda", auxfiles = TRUE, data.set = "default")
Arguments
alpha |
Method-specific tuning or significance parameter. |
maxiter |
Maximum number of iterations. |
precision |
Convergence tolerance. |
hsq |
Heritability parameter. |
out.filename |
Character string retained for compatibility and used as part of the name of an internal file in the R session temporary directory. The temporary file is removed before the function returns. |
auxfiles |
Logical. If |
data.set |
Name of the SelectionTools data set. |
Value
Invisibly returns a numeric matrix containing the lambda results from the RMLA method when auxfiles = TRUE and the compiled calculation succeeds; otherwise NULL.
Estimate or set genomic-selection lambda values using the rmla.02 method
Description
Estimate or set genomic-selection lambda values using the rmla.02 method.
Usage
gs.lambda.rmla.02(alpha = 1, maxiter = 1000, precision = 0.001, hsq = 0.9,
out.filename = "lambda", auxfiles = TRUE,
data.set = "default")
Arguments
alpha |
Method-specific tuning or significance parameter. |
maxiter |
Maximum number of iterations. |
precision |
Convergence tolerance. |
hsq |
Heritability parameter. |
out.filename |
Character string retained for compatibility and used as part of the name of an internal file in the R session temporary directory. The temporary file is removed before the function returns. |
auxfiles |
Logical. If |
data.set |
Name of the SelectionTools data set. |
Value
Invisibly returns a numeric matrix containing the lambda results from the RMLA.02 method when auxfiles = TRUE and the compiled calculation succeeds; otherwise NULL.
Estimate or set genomic-selection lambda values using the rmlc method
Description
Estimate or set genomic-selection lambda values using the rmlc method.
Usage
gs.lambda.rmlc(maxiter = 1000, precision = 0.0001, hsq = 0.9,
out.filename = "lambda", auxfiles = TRUE, data.set = "default")
Arguments
maxiter |
Maximum number of iterations. |
precision |
Convergence tolerance. |
hsq |
Heritability parameter. |
out.filename |
Character string retained for compatibility and used as part of the name of an internal file in the R session temporary directory. The temporary file is removed before the function returns. |
auxfiles |
Logical. If |
data.set |
Name of the SelectionTools data set. |
Value
Invisibly returns a numeric matrix containing the lambda results from the RMLC method when auxfiles = TRUE and the compiled calculation succeeds; otherwise NULL.
Estimate or set genomic-selection lambda values using the rmlr method
Description
Estimate or set genomic-selection lambda values using the rmlr method.
Usage
gs.lambda.rmlr(maxiter = 1000, precision = 0.001, hsq = 0.9,
out.filename = "lambda", auxfiles = TRUE, data.set = "default")
Arguments
maxiter |
Maximum number of iterations. |
precision |
Convergence tolerance. |
hsq |
Heritability parameter. |
out.filename |
Character string retained for compatibility and used as part of the name of an internal file in the R session temporary directory. The temporary file is removed before the function returns. |
auxfiles |
Logical. If |
data.set |
Name of the SelectionTools data set. |
Value
Invisibly returns a numeric matrix containing the lambda results from the RMLR method when auxfiles = TRUE and the compiled calculation succeeds; otherwise NULL.
Estimate or set genomic-selection lambda values using the rmlv method
Description
Estimate or set genomic-selection lambda values using the rmlv method.
Usage
gs.lambda.rmlv(maxiter = 1000, precision = 0.001, out.filename = "lambda",
auxfiles = TRUE, data.set = "default")
Arguments
maxiter |
Maximum number of iterations. |
precision |
Convergence tolerance. |
out.filename |
Character string retained for compatibility and used as part of the name of an internal file in the R session temporary directory. The temporary file is removed before the function returns. |
auxfiles |
Logical. If |
data.set |
Name of the SelectionTools data set. |
Value
Invisibly returns a numeric matrix containing the lambda results from the RMLV method when auxfiles = TRUE and the compiled calculation succeeds; otherwise NULL.
Estimate or set genomic-selection lambda values using the rrblup method
Description
Estimate or set genomic-selection lambda values using the rrblup method.
Usage
gs.lambda.rrblup(maxiter = 1000, precision = 1e-4, hsq = 0.8,
out.filename = "lambda", auxfiles = TRUE,
data.set = "default")
Arguments
maxiter |
Maximum number of iterations. |
precision |
Convergence tolerance. |
hsq |
Heritability parameter. |
out.filename |
Character string retained for compatibility and used as part of the name of an internal file in the R session temporary directory. The temporary file is removed before the function returns. |
auxfiles |
Logical. If |
data.set |
Name of the SelectionTools data set. |
Value
Invisibly returns a numeric matrix containing the lambda results from RR-BLUP when auxfiles = TRUE and the compiled calculation succeeds; otherwise NULL.
Gs mme coeff
Description
Build or return the coefficient matrix for genomic-selection mixed-model equations.
Usage
gs.mme.coeff(out.filename = "mme.coeff", auxfiles = FALSE,
data.set = "default")
Arguments
out.filename |
Character string retained for compatibility and used as part of the name of an internal file in the R session temporary directory. The temporary file is removed before the function returns. |
auxfiles |
Logical. If |
data.set |
Name of the SelectionTools data set. |
Value
Invisibly returns a numeric matrix containing the mixed-model coefficient matrix when auxfiles = TRUE and the compiled calculation succeeds; otherwise NULL.
Gs mme invcoeff
Description
Build or return the inverse coefficient matrix for genomic-selection mixed-model equations.
Usage
gs.mme.invcoeff(out.filename = "inv.mme", auxfiles = TRUE)
Arguments
out.filename |
Character string retained for compatibility and used as part of the name of an internal file in the R session temporary directory. The temporary file is removed before the function returns. |
auxfiles |
Logical. If |
Value
Invisibly returns a numeric matrix containing the inverse mixed-model coefficient matrix when auxfiles = TRUE and the compiled calculation succeeds; otherwise NULL.
Gs mme restcoeff
Description
Build or return the restricted coefficient matrix for genomic-selection mixed-model equations.
Usage
gs.mme.restcoeff(out.filename = "mme.coeff", auxfiles = FALSE)
Arguments
out.filename |
Character string retained for compatibility and used as part of the name of an internal file in the R session temporary directory. The temporary file is removed before the function returns. |
auxfiles |
Logical. If |
Value
Invisibly returns a numeric matrix containing the restricted mixed-model coefficient matrix when auxfiles = TRUE and the compiled calculation succeeds; otherwise NULL.
Gs mme restrhs
Description
Build or return the restricted right-hand side for genomic-selection mixed-model equations.
Usage
gs.mme.restrhs(out.filename = "mme.rhs", auxfiles = FALSE)
Arguments
out.filename |
Character string retained for compatibility and used as part of the name of an internal file in the R session temporary directory. The temporary file is removed before the function returns. |
auxfiles |
Logical. If |
Value
Invisibly returns a numeric matrix containing the restricted mixed-model right-hand side when auxfiles = TRUE and the compiled calculation succeeds; otherwise NULL.
Build or return the right-hand side for genomic-selection mixed-model equations
Description
Build or return the right-hand side for genomic-selection mixed-model equations.
Usage
gs.mme.rhs(out.filename = "mme.rhs", auxfiles = FALSE, data.set = "default")
Arguments
out.filename |
Character string retained for compatibility and used as part of the name of an internal file in the R session temporary directory. The temporary file is removed before the function returns. |
auxfiles |
Logical. If |
data.set |
Name of the SelectionTools data set. |
Value
Invisibly returns a numeric matrix containing the mixed-model right-hand side when auxfiles = TRUE and the compiled calculation succeeds; otherwise NULL.
Gs mme solve
Description
Build or return the mixed-model equation solution for genomic-selection mixed-model equations.
Usage
gs.mme.solve(out.filename = "mme.solution", auxfiles = FALSE,
data.set = "default")
Arguments
out.filename |
Character string retained for compatibility and used as part of the name of an internal file in the R session temporary directory. The temporary file is removed before the function returns. |
auxfiles |
Logical. If |
data.set |
Name of the SelectionTools data set. |
Value
Invisibly returns a numeric matrix containing the mixed-model equation solution when auxfiles = TRUE and the compiled calculation succeeds; otherwise NULL.
Build or solve transformed mixed-model equations (coeff)
Description
Build or solve transformed mixed-model equations (coeff).
Usage
gs.mmet.coeff(out.filename = "mme.coeff", auxfiles = FALSE,
data.set = "default")
Arguments
out.filename |
Character string retained for compatibility and used as part of the name of an internal file in the R session temporary directory. The temporary file is removed before the function returns. |
auxfiles |
Logical. If |
data.set |
Name of the SelectionTools data set. |
Value
Invisibly returns a numeric matrix containing the transformed mixed-model coefficient matrix when auxfiles = TRUE and the compiled calculation succeeds; otherwise NULL.
Build or solve transformed mixed-model equations (coeff.3)
Description
Build or solve transformed mixed-model equations (coeff.3).
Usage
gs.mmet.coeff.3(out.filename = "mme.coeff", auxfiles = FALSE,
data.set = "default")
Arguments
out.filename |
Character string retained for compatibility and used as part of the name of an internal file in the R session temporary directory. The temporary file is removed before the function returns. |
auxfiles |
Logical. If |
data.set |
Name of the SelectionTools data set. |
Value
Invisibly returns a numeric matrix containing the alternative transformed mixed-model coefficient matrix when auxfiles = TRUE and the compiled calculation succeeds; otherwise NULL.
Build or solve transformed mixed-model equations (rhs)
Description
Build or solve transformed mixed-model equations (rhs).
Usage
gs.mmet.rhs(out.filename = "mme.rhs", auxfiles = FALSE, data.set = "default")
Arguments
out.filename |
Character string retained for compatibility and used as part of the name of an internal file in the R session temporary directory. The temporary file is removed before the function returns. |
auxfiles |
Logical. If |
data.set |
Name of the SelectionTools data set. |
Value
Invisibly returns a numeric matrix containing the transformed mixed-model right-hand side when auxfiles = TRUE and the compiled calculation succeeds; otherwise NULL.
Build or solve transformed mixed-model equations (solve)
Description
Build or solve transformed mixed-model equations (solve).
Usage
gs.mmet.solve(out.filename = "mme.solution", auxfiles = FALSE,
data.set = "default")
Arguments
out.filename |
Character string retained for compatibility and used as part of the name of an internal file in the R session temporary directory. The temporary file is removed before the function returns. |
auxfiles |
Logical. If |
data.set |
Name of the SelectionTools data set. |
Value
Invisibly returns a numeric matrix containing the transformed mixed-model solution when auxfiles = TRUE and the compiled calculation succeeds; otherwise NULL.
Set or process all marker effects using the negative-effect convention
Description
Set or process all marker effects using the negative-effect convention.
Usage
gs.neg.effall(data.set ="default")
Arguments
data.set |
Name of the SelectionTools data set. |
Value
On a successful run, an invisible implementation-level list returned by the final timing/information helper; the meaningful result is the side effect of processing marker effects using the negative-effect convention. Validation failures may return NULL.
gs.plot.effects()
Description
Plotting function for the estimated marker effects.
Usage
gs.plot.effects(data.set = "default", absolute = TRUE, alpha = 0.05,
p.adjust.method = "none", pvals = NULL, xlim = NULL,
ylim = NULL, plt = TRUE, col.sp = "blue", pch.sp = 19,
col.sn = "blue", pch.sn = 19, col.ns = "blue", pch.ns = 1,
...)
Arguments
data.set |
character. Name of the data set that is used by the function. |
absolute |
Function argument. |
alpha |
Function argument. |
p.adjust.method |
Function argument. |
pvals |
Function argument. |
xlim |
Function argument. |
ylim |
Function argument. |
plt |
Function argument. |
col.sp |
Function argument. |
pch.sp |
Function argument. |
col.sn |
Function argument. |
pch.sn |
Function argument. |
col.ns |
Function argument. |
pch.ns |
Function argument. |
... |
Additional arguments. |
Value
Invisibly returns the data.frame used for plotting, with columns chrom, pos, name, cpos, and effect, and a pvalue column when p-values are available. Returns NULL if the required effect or map data are unavailable.
gs.plot.model.fit()
Description
Direct genomic values are predicted for those individuals that were used for building the model and estimating marker effects. A scatter plot is created showing the true breeding values (y, phenotypes) and the predicted DGVs (yhat).
Usage
gs.plot.model.fit(training.set ="default", title = "")
Arguments
training.set |
character. Name of the data set to be used for prediction. |
title |
character. Title of the plot. |
Value
No useful return value is intended. The function is called for its side effect of plotting model-fit information; invalid input returns NULL invisibly.
gs.plot.validation()
Description
Creates a graphical output wit two plots. The left plot shows the model fit as produced with the function gs.plot.model.fit. The right plot shows a scatter plot of the predicted DGV and the real phenotypes for a set of individuals that was not used for estimation of marker effects.
Usage
gs.plot.validation(training.set, prediction.set, title="")
Arguments
training.set |
character. Name of the data set that was used for estimation of marker effects. |
prediction.set |
character. Name of the data set that is used for prediction of DGVs within this function. |
title |
character. Title of the plot. |
Details
Within this function, the function gs.estimate.gv is called twice. For the model fit,
the estimation set is both used as estimation and prediction set in gs.estimate.gv.
To validate the model, the function gs.estimate.gv is called again with the
validation set used as prediction set.
The validation set is a set of individuals that was not used for building the model,
but having phenotypes to be correlated to the predicted DGVs.
Value
No useful return value is intended. The function is called for its side effect of plotting validation information; invalid input returns NULL invisibly.
Set or process all marker effects using the positive-effect convention
Description
Set or process all marker effects using the positive-effect convention.
Usage
gs.pos.effall(data.set ="default")
Arguments
data.set |
Name of the SelectionTools data set. |
Value
On a successful run, an invisible implementation-level list returned by the final timing/information helper; the meaningful result is the side effect of processing marker effects using the positive-effect convention. Validation failures may return NULL.
gs.predict.genotypes()
Description
Predicts direct genomic values (DGV) for a given set of individuals in the prediction set. Works only after the marker effects were estimated.
Usage
gs.predict.genotypes(training.set ="default", prediction.set ="default",
out.filename = "breeding.values", auxfiles = TRUE)
Arguments
training.set |
character. Name of the data set that was used for estimating marker effects |
prediction.set |
character. Name of the data set that is used for predicting DGVs |
out.filename |
Character string retained for compatibility and used as part of the name of an internal file in the R session temporary directory. The temporary file is removed before the function returns. |
auxfiles |
Logical. If |
Value
Invisibly returns a data.frame with prediction results when auxfiles = TRUE and prediction succeeds. It contains individual identifiers and predicted genomic values, and also observed phenotypes when those are present in the generated result. Otherwise NULL.
Reset genomic-selection data and settings
Description
Reset genomic-selection data and settings.
Usage
gs.reset()
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of resetting the genomic-selection state.
Restrict marker data using the package genomic-selection interface
Description
Restrict marker data using the package genomic-selection interface.
Usage
gs.restrict.marker.data.01(outfile, traitfile = NULL, NoAll.MAX = 999,
MaMis.MAX = 1, ExHet.MIN = 0, InMis.MAX = 1,
data.set = "default")
Arguments
outfile |
Function argument used by this operation. |
traitfile |
Function argument used by this operation. |
NoAll.MAX |
Function argument used by this operation. |
MaMis.MAX |
Function argument used by this operation. |
ExHet.MIN |
Function argument used by this operation. |
InMis.MAX |
Function argument used by this operation. |
data.set |
Name of the SelectionTools data set. |
Value
On a successful run, an invisible implementation-level list returned by the final timing/information helper; the meaningful result is the side effect of restricting the genomic-selection marker data. Validation failures may return NULL.
gs.return.effects()
Description
Returns the marker effects currently stored in a SelectionTools data set. The underlying routine writes the effects to an intermediate file; the R wrapper reads that file into R and deletes it before returning.
Usage
gs.return.effects(out.filename ="effects", data.set ="default")
Arguments
out.filename |
Character string retained for compatibility and used as part of the name of an internal file in the R session temporary directory. The temporary file is removed before the function returns. |
data.set |
character. Name of the data set that is used by the function. |
Value
A data.frame containing the marker-effect table. If the underlying routine reports an error, NULL is returned invisibly. The intermediate file used to transfer the table is deleted after reading.
Note
This function provides access to effects stored by genomic-effect estimation
functions such as gs.esteff.rr(), gs.esteff.rmlc(), and
gs.esteff.rmlv(). Their auxfiles argument defaults to
FALSE; gs.return.effects() can be used to retrieve the stored
effects independently of that output option.
Return marker p-values for a genomic-selection data set
Description
Return marker p-values for a genomic-selection data set.
Usage
gs.return.pvals(out.filename ="effects", data.set ="default")
Arguments
out.filename |
Character string retained for compatibility and used as part of the name of an internal file in the R session temporary directory. The temporary file is removed before the function returns. |
data.set |
Name of the SelectionTools data set. |
Value
A data.frame containing the marker p-value table. If the underlying routine reports an error, NULL is returned invisibly. The intermediate file used to transfer the table is deleted after reading.
Set all info levels used by the genomic-selection routines
Description
Set all info levels used by the genomic-selection routines.
Usage
gs.set.all.info.levels(level)
Arguments
level |
Function argument used by this operation. |
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of setting all genomic-selection information levels.
Set allele codes used by the genomic-selection routines
Description
Set allele codes used by the genomic-selection routines.
Usage
gs.set.allele.codes(aa = 0, aA = 1, AA = 2, an = 0.5, nn = 1, nA = 1.5,
data.set ="default")
Arguments
aa |
Function argument used by this operation. |
aA |
Function argument used by this operation. |
AA |
Function argument used by this operation. |
an |
Function argument used by this operation. |
nn |
Function argument used by this operation. |
nA |
Function argument used by this operation. |
data.set |
Name of the SelectionTools data set. |
Value
NULL, returned invisibly. The function is called for its side effect of setting the allele codes used by the genomic-selection routines.
Set effects used by the genomic-selection routines
Description
Set effects used by the genomic-selection routines.
Usage
gs.set.effects(eff, filename ="seteffects", data.set ="default")
Arguments
eff |
Function argument used by this operation. |
filename |
Character string retained for compatibility and used as part of the name of the internal session-temporary file used to pass the effects to compiled code. The file is removed before the function returns. |
data.set |
Name of the SelectionTools data set. |
Value
On a successful run, an invisible implementation-level list returned by the final timing/information helper; the meaningful result is the side effect of setting marker effects in the genomic-selection data set. Validation failures may return NULL.
Set info level used by the genomic-selection routines
Description
Set info level used by the genomic-selection routines.
Usage
gs.set.info.level(level, data.set ="default")
Arguments
level |
Function argument used by this operation. |
data.set |
Name of the SelectionTools data set. |
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of setting the genomic-selection information level for the data set.
Set lambda used by the genomic-selection routines
Description
Set lambda used by the genomic-selection routines.
Usage
gs.set.lambda(lambda, out.filename = "lambda", auxfiles = FALSE,
data.set = "default")
Arguments
lambda |
Regularization parameter. |
out.filename |
Character string retained for compatibility and used as part of the name of an internal file in the R session temporary directory. The temporary file is removed before the function returns. |
auxfiles |
Logical. If |
data.set |
Name of the SelectionTools data set. |
Value
Invisibly returns a numeric matrix containing the lambda values produced while setting lambda when auxfiles = TRUE and the compiled calculation succeeds; otherwise NULL.
Set num threads used by the genomic-selection routines
Description
Set num threads used by the genomic-selection routines.
Usage
gs.set.num.threads(active)
Arguments
active |
Function argument used by this operation. |
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of setting the number of genomic-selection computational threads.
Set performance data used by the genomic-selection routines
Description
Set performance data used by the genomic-selection routines.
Usage
gs.set.performance.data(y, data.set = "default")
Arguments
y |
Response vector. |
data.set |
Name of the SelectionTools data set. |
Value
An invisible integer scalar containing the status value returned by the compiled performance-data setter.
Perform single-marker analysis for a genomic-selection data set
Description
Perform single-marker analysis for a genomic-selection data set.
Usage
gs.single.marker.aov(alpha = 1, out.filename = "sm.reg", auxfiles = TRUE,
data.set = "default")
Arguments
alpha |
Method-specific tuning or significance parameter. |
out.filename |
Character string retained for compatibility and used as part of the name of an internal file in the R session temporary directory. The temporary file is removed before the function returns. |
auxfiles |
Logical. If |
data.set |
Name of the SelectionTools data set. |
Value
Invisibly returns a numeric matrix containing the single-marker ANOVA results when auxfiles = TRUE and the compiled calculation succeeds; otherwise NULL.
Perform single-marker analysis for a genomic-selection data set
Description
Perform single-marker analysis for a genomic-selection data set.
Usage
gs.single.marker.reg(out.filename = "sm.reg", auxfiles = TRUE,
data.set = "default")
Arguments
out.filename |
Character string retained for compatibility and used as part of the name of an internal file in the R session temporary directory. The temporary file is removed before the function returns. |
auxfiles |
Logical. If |
data.set |
Name of the SelectionTools data set. |
Value
Invisibly returns a numeric matrix containing the single-marker regression results when auxfiles = TRUE and the compiled calculation succeeds; otherwise NULL.
Start or stop the genomic-selection timing helper
Description
Start or stop the genomic-selection timing helper.
Usage
gs.start.timer(depth=1)
Arguments
depth |
Function argument used by this operation. |
Value
An invisible implementation-level list returned by the SelectionTools timing helper. The function is called for the side effect of starting timing.
Start or stop the genomic-selection timing helper
Description
Start or stop the genomic-selection timing helper.
Usage
gs.stop.timer(depth=1, info.level=0)
Arguments
depth |
Function argument used by this operation. |
info.level |
Function argument used by this operation. |
Value
An invisible implementation-level list returned by the SelectionTools information helper after reporting elapsed time. The function is called for the timing/reporting side effect.
Test marker or model parameters used by genomic-selection routines
Description
Test marker or model parameters used by genomic-selection routines.
Usage
gs.test.mp(n = 0)
Arguments
n |
Number of items or repetitions. |
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of running the compiled marker/model-parameter test.
Test estimated marker effects using the indicated method
Description
Test estimated marker effects using the indicated method.
Usage
gs.testeff.lsq(nperm = 0, out.filename = "test.lsq", auxfiles = TRUE,
data.set ="default")
Arguments
nperm |
Function argument used by this operation. |
out.filename |
Character string retained for compatibility and used as part of the name of an internal file in the R session temporary directory. The temporary file is removed before the function returns. |
auxfiles |
Logical. If |
data.set |
Name of the SelectionTools data set. |
Value
Invisibly returns a numeric matrix containing the LSQ effect-test results when auxfiles = TRUE and the compiled calculation succeeds; otherwise NULL.
Test estimated marker effects using the indicated method
Description
Test estimated marker effects using the indicated method.
Usage
gs.testeff.rmlc(nperm = 1000, maxiter = 1000, precision = 0.001, hsq = 0.9,
out.filename = "test.rmlc", auxfiles = TRUE,
data.set ="default")
Arguments
nperm |
Function argument used by this operation. |
maxiter |
Maximum number of iterations. |
precision |
Convergence tolerance. |
hsq |
Heritability parameter. |
out.filename |
Character string retained for compatibility and used as part of the name of an internal file in the R session temporary directory. The temporary file is removed before the function returns. |
auxfiles |
Logical. If |
data.set |
Name of the SelectionTools data set. |
Value
Invisibly returns a data.frame containing the RMLC effect-test results when auxfiles = TRUE and the test succeeds; otherwise NULL.
Test estimated marker effects using the indicated method
Description
Test estimated marker effects using the indicated method.
Usage
gs.testeff.rmlv(nperm = 1000, maxiter = 1000, precision = 0.001, hsq = 0.9,
out.filename = "test.rmlv", auxfiles = TRUE,
data.set ="default")
Arguments
nperm |
Function argument used by this operation. |
maxiter |
Maximum number of iterations. |
precision |
Convergence tolerance. |
hsq |
Heritability parameter. |
out.filename |
Character string retained for compatibility and used as part of the name of an internal file in the R session temporary directory. The temporary file is removed before the function returns. |
auxfiles |
Logical. If |
data.set |
Name of the SelectionTools data set. |
Value
Invisibly returns a numeric matrix containing the RMLV effect-test results when auxfiles = TRUE and the compiled calculation succeeds; otherwise NULL.
Estimate variance components for ridge-regression BLUP
Description
Estimate variance components for ridge-regression BLUP.
Usage
gs.vc.rrblup(maxiter = 1000, precision = 1e-4, hsq = 0.8,
out.filename = "lambda", auxfiles = TRUE, data.set = "default")
Arguments
maxiter |
Maximum number of iterations. |
precision |
Convergence tolerance. |
hsq |
Heritability parameter. |
out.filename |
Character string retained for compatibility and used as part of the name of an internal file in the R session temporary directory. The temporary file is removed before the function returns. |
auxfiles |
Logical. If |
data.set |
Name of the SelectionTools data set. |
Value
Invisibly returns a numeric matrix containing the variance-component results from RR-BLUP when auxfiles = TRUE and the compiled calculation succeeds; otherwise NULL.
Calculate weights or statistics used by the genomic-selection ANOVA method
Description
Calculate weights or statistics used by the genomic-selection ANOVA method.
Usage
gs.w.aov(alpha = 1, out.filename = "W", auxfiles = TRUE, data.set = "default")
Arguments
alpha |
Method-specific tuning or significance parameter. |
out.filename |
Character string retained for compatibility and used as part of the name of an internal file in the R session temporary directory. The temporary file is removed before the function returns. |
auxfiles |
Logical. If |
data.set |
Name of the SelectionTools data set. |
Value
Invisibly returns a numeric matrix containing the ANOVA weights/statistics when auxfiles = TRUE and the compiled calculation succeeds; otherwise NULL.
Write Marker Effects for the Simulation Backend
Description
Writes estimated marker effects in the effect-file format used by the simulation backend.
Usage
gs.write.pseff(e, beta0 = NA, file)
Arguments
e |
A marker-effect table, typically returned by
|
beta0 |
Optional general mean on the genomic-selection scale. If
|
file |
Character string naming the effect file to be written. This argument is required because the function writes a persistent file. |
Details
The genomic-selection and simulation parts of SelectionTools use different origins for additive marker coding.
With the default allele codes used by the genomic-selection routines,
gs.set.allele.codes() assigns the additive dosage codes 0, 1, and 2
to the genotypes with zero, one, and two copies of the effect allele,
respectively. For marker j with estimated effect u_j, the
marker contribution is therefore 0, u_j, or 2u_j.
The simulation effect-file representation stores an additive effect per
allele. The effect allele is written with effect +u_j/2 and the
complementary allele with effect -u_j/2. The corresponding diploid
marker contributions are therefore -u_j, 0, and +u_j.
This is equivalent to changing the marker dosage from z_{ij} to
z^*_{ij}=z_{ij}-1.
On the genomic-selection scale,
\hat{g}_i=\hat{\mu}_{012}+\sum_{j=1}^{m}z_{ij}\hat{u}_j.
Using z_{ij}=z^*_{ij}+1 gives
\hat{g}_i=
\left(\hat{\mu}_{012}+\sum_{j=1}^{m}\hat{u}_j\right)
+\sum_{j=1}^{m}z^*_{ij}\hat{u}_j.
Consequently, the first line written to the effect file is
\hat{\mu}_{\mathrm{sim}}=
\hat{\mu}_{012}+\sum_{j=1}^{m}\hat{u}_j.
If beta0 is supplied, it replaces \hat{\mu}_{012} before this
coding correction is applied.
This is the coding-shift correction described by Stranden and
Christensen (2011), Equation (2). The shift is distinct from
allele-frequency centering. For coding
z^c_{ij}=z_{ij}-2p_j, the corresponding mean shift is
2\sum_j p_j u_j; the simulation effect-file coding instead uses
z^*_{ij}=z_{ij}-1.
The exact equivalence described above applies to the default allele codes.
gs.set.allele.codes() permits power users to define other codes; an
arbitrary custom coding is not in general represented by the fixed
-u_j/2, +u_j/2 simulation allele effects.
Value
NULL, returned invisibly by the file-writing operation. The function is called for its side effect of writing marker effects for the simulation backend.
References
Stranden I, Christensen OF (2011) Allele coding in genomic evaluation. Genetics Selection Evolution 43:25. See Equation (2).
See Also
gs.return.effects, gs.set.allele.codes,
st.set.sim.ef, load.effmap
homozygote()
Description
Produces a dataframe of a population whith NoInd individuals which carry at all loci the same allele (allele). This dataframe can be used to initialise a population.
Usage
homozygote(allele, NoInd = 1)
Arguments
allele |
Allele which is carried at all loci |
NoInd |
Number of individuals to be generated |
Value
A data.frame with columns ind, chrom, hom, pos, and all, representing the requested homozygous genome.
Note
See init.population for an example
Construct or evaluate introgression-line libraries
Description
Construct or evaluate introgression-line libraries.
Usage
il.eval.library(rp.allele = 1, dp.allele = 2, chrom = 0, lower = 0, upper = 0,
data.set = "default")
Arguments
rp.allele |
Allele-related parameter. |
dp.allele |
Allele-related parameter. |
chrom |
Chromosome identifier. |
lower |
Function argument used by this operation. |
upper |
Function argument used by this operation. |
data.set |
Name of the SelectionTools data set. |
Value
A named numeric vector summarizing the evaluated introgression-line library. It contains cov, dep, seglib, dpg, ndps, and ldps; when target regions are supplied it additionally contains tr.dpg, nt.rpg, nt.ndps, and nt.ldps. Returns NULL on validation failure.
Construct or evaluate introgression-line libraries
Description
Construct or evaluate introgression-line libraries.
Usage
il.eval.lines(allele = 1, chrom = 0, lower = 0, upper = 0, exclude = 0,
data.set = "default")
Arguments
allele |
Allele specification. |
chrom |
Chromosome identifier. |
lower |
Function argument used by this operation. |
upper |
Function argument used by this operation. |
exclude |
Function argument used by this operation. |
data.set |
Name of the SelectionTools data set. |
Value
A data.frame with one row per evaluated line and columns name, nseg, lseg, abs, tot, and rel, summarizing introgressed segments and their lengths/proportions. Returns NULL on validation failure.
Construct or evaluate introgression-line libraries
Description
Construct or evaluate introgression-line libraries.
Usage
il.ideal.library(n.c = 5, l.c = 100, s.l = 20, data.set = "default")
Arguments
n.c |
Number or size used by the operation. |
l.c |
Function argument used by this operation. |
s.l |
Function argument used by this operation. |
data.set |
Name of the SelectionTools data set. |
Value
An integer cleanup status returned by the final unlink() call after construction of the ideal introgression-line library. The main result is the population/library created by the function.
Construct or evaluate introgression-line libraries
Description
Construct or evaluate introgression-line libraries.
Usage
il.overlapping.library(n.c = 5, l.c = 100, s.l = 20, data.set = "default")
Arguments
n.c |
Number or size used by the operation. |
l.c |
Function argument used by this operation. |
s.l |
Function argument used by this operation. |
data.set |
Name of the SelectionTools data set. |
Value
An integer cleanup status returned by the final unlink() call after construction of the overlapping introgression-line library. The main result is the population/library created by the function.
Print or handle an information message
Description
Print or handle an information message.
Usage
info(lev, msg)
Arguments
lev |
Information or verbosity level. |
msg |
Message text. |
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of issuing the requested information message.
Emit information-level-controlled status output
Description
Emit a status message when the requested information level is enabled.
The message is signaled with message() and can therefore also be
suppressed with suppressMessages().
Usage
info.cat(lev, msg)
Arguments
lev |
Information or verbosity level. |
msg |
Message text. |
Value
NULL. The function is called for its side effect of conditionally signaling a status message according to the current information level.
init.population()
Description
Creates a simulation population from the internal chromosome-segment data-frame representation.
Usage
init.population(name, data, delete=TRUE)
Arguments
name |
Character string naming the simulation population to initialize. |
data |
Data frame with five columns representing individual, chromosome, homologue, segment position, and allele. |
delete |
Logical. If true, remove an existing population with the same name before initialization. |
Details
The input data frame must have exactly five columns named or interpretable as individual number, chromosome number, homologue number, position, and allele. Each row describes the start of a chromosome segment carrying the specified allele. The data must be ordered consistently with the simulation representation.
The number of input rows must be positive. Chromosome and homologue identifiers must be compatible with the currently defined simulation genome. The routine constructs the raw chromosome representation; marker genotypes and evaluated genetic values are derived data and are not implied merely by initialization.
When delete=TRUE, an existing population of the same name is removed before initialization. If creation fails while chromosome segments are being inserted, the incomplete new population is removed rather than retained.
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of initializing the simulation population from the supplied data.
Linkage Map for Marker-Assisted Backcrossing Examples
Description
A linkage-map data set used in the marker-assisted backcrossing examples based on Frisch et al. (1999). This version represents the map with approximately 80 marker loci used for background selection.
Usage
data(lib00.map1)
Format
An R data object containing linkage-map records in the format used by SelectionTools linkage-map input routines.
Source
Data used for the examples based on Frisch et al. (1999), Crop Science 39:1295–1301.
Reduced Linkage Map for Marker-Assisted Backcrossing Examples
Description
A reduced linkage-map data set used in the marker-assisted backcrossing examples based on Frisch et al. (1999). This version represents the reduced map with approximately 40 marker loci used to compare marker density for background selection.
Usage
data(lib00.map1a)
Format
An R data object containing linkage-map records in the format used by SelectionTools linkage-map input routines.
Source
Data used for the examples based on Frisch et al. (1999), Crop Science 39:1295–1301.
Linkage Map for Three-Stage Marker-Assisted Backcrossing
Description
A linkage-map data set used in the three-stage marker-assisted backcrossing examples based on Frisch et al. (1999). The example uses target, flanking and background-marker information for selection.
Usage
data(lib00.map2)
Format
An R data object containing linkage-map records in the format used by SelectionTools linkage-map input routines.
Source
Data used for the examples based on Frisch et al. (1999), Crop Science 39:1295–1301.
Calculate or summarize linkage drag
Description
Calculate or summarize linkage drag.
Usage
linkage.drag(pops, allele, chrom, pos)
Arguments
pops |
Population name or names. |
allele |
Allele specification. |
chrom |
Chromosome identifier. |
pos |
Position on the chromosome. |
Value
A numeric matrix with one row per population and columns Obs, Mean, SDev, Min, Q10, Med, and Max, summarizing linkage-drag segment lengths.
Get a linkage map used by the simulation routines
Description
Returns the currently active simulation linkage map.
Usage
linkage.map.get()
Details
The result contains the map points used by the simulation backend, including chromosome, position, locus name, and class, in the current simulation map order.
Value
A data.frame with columns chrom, pos, name, and class, describing the active simulation linkage map.
Load a linkage map used by the simulation routines
Description
Loads and prepares the linkage map used by the simulation backend.
Usage
linkage.map.load(file, disperse=TRUE, file.disperse=NA, disperse.factor=100)
Arguments
file |
Character string naming the simulation map file. |
disperse |
Logical flag controlling dispersion of coincident map positions. |
file.disperse |
Optional file-related dispersion specification used by the existing map loader. |
disperse.factor |
Numeric factor controlling the dispersion scale used by the existing loader. |
Details
The simulation map supplies chromosome, position, locus-name, and class information used by population genotyping and recombination. The loader prepares map points in a valid access order and can disperse coincident positions when requested. The active simulation map is a simulation-wide object and is separate from a SelectionTools marker data set.
Value
An invisible list with a retval status component. Invalid input returns list(retval = -2L) invisibly; otherwise the list is returned by the compiled map loader. The main effect is loading the linkage map.
Save a linkage map used by the simulation routines
Description
Writes a simulation linkage map to a text file.
Usage
linkage.map.save(file, dta.map=linkage.map.get())
Arguments
file |
Character string naming the output map file. |
dta.map |
Map data to write; by default the currently active simulation linkage map. |
Details
By default the currently active simulation linkage map is obtained and written. A supplied dta.map can instead be written. Saving does not alter the active map or any simulation population.
Value
NULL, returned invisibly by write.table(). The function is called for its side effect of writing the linkage map.
list.effects()
Description
Rerturns all defined effects and their weights.
Usage
list.effects()
Value
A data.frame with columns effect and weight, listing the currently defined effects and their weights.
Note
See define.effects for for an example
list.populations
Description
Returns the names and active sizes of the currently registered simulation populations.
Usage
list.populations()
Details
The result reflects the simulation population registry and is independent of SelectionTools marker-data-set names.
Value
Returns the same data.frame as population.list(), with columns PopName and count.
load.effmap()
Description
Loads an effect file
Usage
load.effmap(name, file = NA)
Arguments
name |
Name of the effect |
file |
File that contains the effect description |
Details
Effect files are text files can be generated either by hand with
an editor or with the command generate.effect.file().
An effect file consists of a value, which is assigned to each individual
of a population (the population mean) and a list of effects which are
added to the population mean if the certain allelic combinations occur
in an individual. The order of the effects is arbitrary, but the
population mean must occur first in the map file.
Mean;
Structure of an effect file:
Locus Allele [Locus Allele ...] Effect;
[,...]
| Mean | The population mean |
| Classname | A class of Loci to which effects are assigned to |
| Allele | The alleles to which the effects are assigned |
| [Allele] | Optional. Is used to define dominance effects epistatic effects |
| [,...] | An arbitrary number of effects can be defined seperated by commas |
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of loading the requested effect map.
Note
See define.effects for for an example
Load an internal effect map
Description
Load an internal effect map.
Usage
load.internal.effmap(name, spec, weight)
Arguments
name |
Function argument used by this operation. |
spec |
Function argument used by this operation. |
weight |
Weight value. |
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of loading the requested internal effect map.
load.linkage.map()
Description
Loads the simulation linkage map from a text file.
Usage
load.linkage.map(file, disperse=TRUE, file.disperse=NA, disperse.factor=100)
Arguments
file |
Character string naming the simulation linkage-map input file. |
disperse |
Logical. If true, coincident locus positions are deterministically separated during map preparation. |
file.disperse |
Optional character string naming a file to which the dispersed map is written. An omitted value does not create an additional file. |
disperse.factor |
Finite numeric factor controlling the spacing used when coincident positions are dispersed. |
Details
The simulation linkage map is distinct from the marker-data map managed by st.read.map. It defines the loci available to simulation-population genotyping, recombination, and evaluation.
Map records contain chromosome, position, locus name, and class. The loader prepares the loci in chromosome/position order and can disperse coincident map points according to its arguments. Only one simulation linkage map is active at a time.
The function linkage.map.load provides the same linkage-map loading operation.
Value
Returns the same invisible status list as linkage.map.load(), including its retval component.
mab.Examples
Description
Examples for the optimize-mab routines
Value
No return value. This is a documentation page describing examples for the marker-assisted backcrossing routines.
See Also
mab.simulate, mab.compare,
mab.load.data, mab.tabulate,
mab.Input.files
Structure of input files for the optimize-mab routines.
Description
Structure of the files describing linkage maps and crossing parents.
Files must be placed in the directory st.input.dir.
Details
A linkage-map file contains chromosome number, position in cM, and marker name. The old package documentation gave, for example:
1 0 umc94 1 0002 bnl805 1 0067 umc76 1 0070 umc137 2 0 bnl845 2 0012 umc53
For a homozygous crossing parent, the first line is the genotype name and each following line contains a marker name and one allele, for example:
dh.2010.12345 s1e2855 ade s1e3177 cyt s1e2083 ade s1e4552 gua
For a heterozygous crossing parent, the first line is again the genotype name and each following line contains a marker name and two alleles, for example:
bc1i3 umc94 1 8 bnl805 1 8 umc76 8 8 umc137 8 8
Value
No return value. This is a documentation page describing the input-file formats used by the marker-assisted backcrossing routines.
See Also
Run and Compare a Series of Marker-Assisted Backcrossing Simulations
Description
Runs up to nine marker-assisted backcrossing simulation scenarios by calling
mab.simulate() for each active scenario. Scenario A provides the base
settings; scenarios B through I can override individual settings.
Usage
mab.compare(
a.simulation.name,
a.simulation.run,
a.repetitions = 1000,
a.recurrent.parent = "H",
a.donor.parent = "H",
a.linkage.map,
a.target.loci = NULL,
a.flanking.loci = NULL,
a.recipient.loci = NULL,
a.gen.type = NULL,
a.population.size = NULL,
a.sel.strategy = NULL,
a.no.selected = NULL,
a.no.preselected = NULL,
a.reg.chr = NULL,
a.reg.begin = NULL,
a.reg.end = NULL,
a.missing.allele = "9",
a.success.factor = 1,
b.simulation.name = NULL,
c.simulation.name = NULL,
d.simulation.name = NULL,
b.simulation.run = NULL,
c.simulation.run = NULL,
d.simulation.run = NULL,
b.repetitions = NULL,
c.repetitions = NULL,
d.repetitions = NULL,
b.recurrent.parent = NULL,
c.recurrent.parent = NULL,
d.recurrent.parent = NULL,
b.donor.parent = NULL,
c.donor.parent = NULL,
d.donor.parent = NULL,
b.linkage.map = NULL,
c.linkage.map = NULL,
d.linkage.map = NULL,
b.target.loci = NULL,
c.target.loci = NULL,
d.target.loci = NULL,
b.flanking.loci = NULL,
c.flanking.loci = NULL,
d.flanking.loci = NULL,
b.recipient.loci = NULL,
c.recipient.loci = NULL,
d.recipient.loci = NULL,
b.gen.type = NULL,
c.gen.type = NULL,
d.gen.type = NULL,
b.population.size = NULL,
c.population.size = NULL,
d.population.size = NULL,
b.sel.strategy = NULL,
c.sel.strategy = NULL,
d.sel.strategy = NULL,
b.no.selected = NULL,
c.no.selected = NULL,
d.no.selected = NULL,
b.no.preselected = NULL,
c.no.preselected = NULL,
d.no.preselected = NULL,
b.reg.chr = NULL,
c.reg.chr = NULL,
d.reg.chr = NULL,
b.reg.begin = NULL,
c.reg.begin = NULL,
d.reg.begin = NULL,
b.reg.end = NULL,
c.reg.end = NULL,
d.reg.end = NULL,
b.missing.allele = NULL,
c.missing.allele = NULL,
d.missing.allele = NULL,
b.success.factor = NULL,
c.success.factor = NULL,
d.success.factor = NULL,
e.simulation.name = NULL,
f.simulation.name = NULL,
g.simulation.name = NULL,
e.simulation.run = NULL,
f.simulation.run = NULL,
g.simulation.run = NULL,
e.repetitions = NULL,
f.repetitions = NULL,
g.repetitions = NULL,
e.recurrent.parent = NULL,
f.recurrent.parent = NULL,
g.recurrent.parent = NULL,
e.donor.parent = NULL,
f.donor.parent = NULL,
g.donor.parent = NULL,
e.linkage.map = NULL,
f.linkage.map = NULL,
g.linkage.map = NULL,
e.target.loci = NULL,
f.target.loci = NULL,
g.target.loci = NULL,
e.flanking.loci = NULL,
f.flanking.loci = NULL,
g.flanking.loci = NULL,
e.recipient.loci = NULL,
f.recipient.loci = NULL,
g.recipient.loci = NULL,
e.gen.type = NULL,
f.gen.type = NULL,
g.gen.type = NULL,
e.population.size = NULL,
f.population.size = NULL,
g.population.size = NULL,
e.sel.strategy = NULL,
f.sel.strategy = NULL,
g.sel.strategy = NULL,
e.no.selected = NULL,
f.no.selected = NULL,
g.no.selected = NULL,
e.no.preselected = NULL,
f.no.preselected = NULL,
g.no.preselected = NULL,
e.reg.chr = NULL,
f.reg.chr = NULL,
g.reg.chr = NULL,
e.reg.begin = NULL,
f.reg.begin = NULL,
g.reg.begin = NULL,
e.reg.end = NULL,
f.reg.end = NULL,
g.reg.end = NULL,
e.missing.allele = NULL,
f.missing.allele = NULL,
g.missing.allele = NULL,
e.success.factor = NULL,
f.success.factor = NULL,
g.success.factor = NULL,
h.simulation.name = NULL,
i.simulation.name = NULL,
h.simulation.run = NULL,
i.simulation.run = NULL,
h.repetitions = NULL,
i.repetitions = NULL,
h.recurrent.parent = NULL,
i.recurrent.parent = NULL,
h.donor.parent = NULL,
i.donor.parent = NULL,
h.linkage.map = NULL,
i.linkage.map = NULL,
h.target.loci = NULL,
i.target.loci = NULL,
h.flanking.loci = NULL,
i.flanking.loci = NULL,
h.recipient.loci = NULL,
i.recipient.loci = NULL,
h.gen.type = NULL,
i.gen.type = NULL,
h.population.size = NULL,
i.population.size = NULL,
h.sel.strategy = NULL,
i.sel.strategy = NULL,
h.no.selected = NULL,
i.no.selected = NULL,
h.no.preselected = NULL,
i.no.preselected = NULL,
h.reg.chr = NULL,
i.reg.chr = NULL,
h.reg.begin = NULL,
i.reg.begin = NULL,
h.reg.end = NULL,
i.reg.end = NULL,
h.missing.allele = NULL,
i.missing.allele = NULL,
h.success.factor = NULL,
i.success.factor = NULL,
result.files
)
Arguments
a.simulation.name, b.simulation.name, c.simulation.name, d.simulation.name, e.simulation.name, f.simulation.name, g.simulation.name, h.simulation.name, i.simulation.name |
Simulation name for each scenario. Scenario A is required. For scenarios B through I, a non- |
a.simulation.run, b.simulation.run, c.simulation.run, d.simulation.run, e.simulation.run, f.simulation.run, g.simulation.run, h.simulation.run, i.simulation.run |
Simulation-run identifier passed to |
a.repetitions, b.repetitions, c.repetitions, d.repetitions, e.repetitions, f.repetitions, g.repetitions, h.repetitions, i.repetitions |
Number of simulation repetitions. Scenario A defaults to 1000. For active scenarios B through I, |
a.recurrent.parent, b.recurrent.parent, c.recurrent.parent, d.recurrent.parent, e.recurrent.parent, f.recurrent.parent, g.recurrent.parent, h.recurrent.parent, i.recurrent.parent |
Description of the recurrent parent passed to |
a.donor.parent, b.donor.parent, c.donor.parent, d.donor.parent, e.donor.parent, f.donor.parent, g.donor.parent, h.donor.parent, i.donor.parent |
Description of the donor parent passed to |
a.linkage.map, b.linkage.map, c.linkage.map, d.linkage.map, e.linkage.map, f.linkage.map, g.linkage.map, h.linkage.map, i.linkage.map |
Name of the linkage-map input file passed to |
a.target.loci, b.target.loci, c.target.loci, d.target.loci, e.target.loci, f.target.loci, g.target.loci, h.target.loci, i.target.loci |
Character vector of target-locus names passed to |
a.flanking.loci, b.flanking.loci, c.flanking.loci, d.flanking.loci, e.flanking.loci, f.flanking.loci, g.flanking.loci, h.flanking.loci, i.flanking.loci |
Character vector of flanking-locus names used for recurrent-parent preselection, or |
a.recipient.loci, b.recipient.loci, c.recipient.loci, d.recipient.loci, e.recipient.loci, f.recipient.loci, g.recipient.loci, h.recipient.loci, i.recipient.loci |
Character vector of loci used for preselection of recurrent-parent alleles, or |
a.gen.type, b.gen.type, c.gen.type, d.gen.type, e.gen.type, f.gen.type, g.gen.type, h.gen.type, i.gen.type |
Character vector describing the generation types passed to |
a.population.size, b.population.size, c.population.size, d.population.size, e.population.size, f.population.size, g.population.size, h.population.size, i.population.size |
Numeric vector of population sizes for the generations of the backcrossing program, or |
a.sel.strategy, b.sel.strategy, c.sel.strategy, d.sel.strategy, e.sel.strategy, f.sel.strategy, g.sel.strategy, h.sel.strategy, i.sel.strategy |
Character vector of selection-strategy codes passed to |
a.no.selected, b.no.selected, c.no.selected, d.no.selected, e.no.selected, f.no.selected, g.no.selected, h.no.selected, i.no.selected |
Numeric vector giving the number of selected individuals by generation, or |
a.no.preselected, b.no.preselected, c.no.preselected, d.no.preselected, e.no.preselected, f.no.preselected, g.no.preselected, h.no.preselected, i.no.preselected |
Numeric vector giving the minimum number of individuals retained by the flanking-marker preselection step, or |
a.reg.chr, b.reg.chr, c.reg.chr, d.reg.chr, e.reg.chr, f.reg.chr, g.reg.chr, h.reg.chr, i.reg.chr |
Chromosome numbers for regions evaluated separately for recurrent-parent genome content, or |
a.reg.begin, b.reg.begin, c.reg.begin, d.reg.begin, e.reg.begin, f.reg.begin, g.reg.begin, h.reg.begin, i.reg.begin |
Beginning positions of the regions specified by |
a.reg.end, b.reg.end, c.reg.end, d.reg.end, e.reg.end, f.reg.end, g.reg.end, h.reg.end, i.reg.end |
End positions of the regions specified by |
a.missing.allele, b.missing.allele, c.missing.allele, d.missing.allele, e.missing.allele, f.missing.allele, g.missing.allele, h.missing.allele, i.missing.allele |
Character string used by |
a.success.factor, b.success.factor, c.success.factor, d.success.factor, e.success.factor, f.success.factor, g.success.factor, h.success.factor, i.success.factor |
Success factor passed to |
result.files |
Character vector containing one explicit result-file name for each active scenario, in scenario order A through I. The names must be non-empty and unique. Relative names are resolved against |
Details
The scenario prefixes a. through i. identify up to nine
simulation definitions. Scenario A is always defined by
a.simulation.name. A scenario B through I is run only when its
simulation.name argument is non-NULL.
For every setting other than simulation.name, an active scenario B
through I inherits the corresponding scenario-A value when its own argument is
NULL; a non-NULL scenario-specific value overrides the
scenario-A value. The resulting settings are passed directly to
mab.simulate(). See mab.simulate() for the detailed meaning of
the generation and selection settings and for the structure of the simulation
results. The file name for every active scenario is taken from
result.files; no result-file name is generated automatically.
Value
NULL, returned invisibly as the value of the final for loop.
The function is called for its side effect of running the active simulation
scenarios; each call to mab.simulate() saves its result in the
corresponding explicitly supplied file from result.files.
See Also
mab.simulate, mab.load.data,
mab.tabulate, mab.Input.files,
mab.Examples
Helper used by the marker-assisted backcross simulation routines
Description
Resolves MAB file names against a configured directory while preserving explicit absolute paths.
Usage
mab.df(d, f)
Arguments
d |
Directory used for a relative file name. |
f |
Non-empty file name or path. Absolute paths are returned unchanged. |
Value
A character scalar. If f is absolute it is returned unchanged. If f is relative and d is non-empty, the two are combined with file.path(); otherwise f is returned unchanged.
Loads the data of a simulation
Description
Loads a simulation result from the explicitly supplied result.file
for further analysis or printing. Relative file names are resolved against
st.data.dir; absolute paths are used directly.
Usage
mab.load.data(simulation.name, simulation.run, result.file)
Arguments
simulation.name |
Simulation-name identifier retained for compatibility with the MAB workflow. It does not determine the file name. |
simulation.run |
Simulation-run identifier retained for compatibility with the MAB workflow. It does not determine the file name. |
result.file |
Required name of the result file. Relative names are resolved against |
Value
The object outp loaded from the requested simulation result file. For files created by mab.simulate(), this is the named simulation-result list described there.
See Also
mab.simulate, mab.tabulate,
mab.Examples
Helper used by the marker-assisted backcross simulation routines
Description
Writes one marker-data file for each individual in a population. Output file
names use the explicitly supplied file.prefix.
Usage
mab.save.inds(pop, pname, file.prefix)
Arguments
pop |
Population name or population specification. |
pname |
Prefix used for the individual names written inside the marker-data files. |
file.prefix |
Required prefix for the output file names. Each file receives a three-digit individual number and the extension |
Value
NULL. The function is called for its side effect of writing one marker-data file for each individual in the supplied population.
Simulates a marker-assisted backcrossing program
Description
Simulates one marker-assisted backcrossing program. The simulation result
is stored in the explicitly supplied result.file. A relative file
name is resolved against st.data.dir; an absolute path is used as
supplied.
Usage
mab.simulate(simulation.name, simulation.run, repetitions = 1000,
recurrent.parent = "H", donor.parent = "H", linkage.map,
target.loci = NULL, flanking.loci = NULL, recipient.loci = NULL,
gen.type = NULL, population.size = NULL, sel.strategy = NULL,
no.selected = NULL, no.preselected = NULL, reg.chr = NULL,
reg.begin = NULL, reg.end = NULL, missing.allele = "9",
result.file, success.factor = 1, recode.infiles = FALSE,
recode.file = NULL)
Arguments
simulation.name |
Name of the simulation stored in the returned result object. The argument is required and has no default value. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
simulation.run |
Name of the simulation run. Usually a single character. This can be
used to run a certain simulation for several times. For example, in a
test run with 1000 repetitions, the character | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
repetitions |
Number of repetitions of the simulation. The default is 1000. For final analyses, a larger number of repetitions may be used. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
recurrent.parent |
Description of the recurrent parent. If | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
donor.parent |
Description of the donor parent. See | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
linkage.map |
Name of the file, in which the linkage map is described. See the
section 'Input Files' for a description of the file structure.The file
must be located in the directory | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
target.loci |
Character vector with names of target loci. Selection for the donor
allele at | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
flanking.loci |
Character vector with names of loci, which are used to carry out
pre selection for the recurrent parent allele before genome wide
background selection is carried out. Selection for the recipient allele
at | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
recipient.loci |
Character vector, which contains the names of loci wich are used to carry
out preselection for the recipient alleles. Selection for the recipient
allele at | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
gen.type |
A character vector consisting of a description of the generation types of the backcrossing program. The following generation types are implemented:
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
population.size |
Numeric vector containing the population sizes of the
generations in the backcrossing program. The formal default is
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
sel.strategy |
Character vector consisting of the selection strategies for the different generations of a backcrossing program. The following selection strategies are implemented. SI: selection index.
Selection strategies Counting marker data points:
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
no.selected |
Numerical vector specifying the number of selected individuals in the backcross generations. Note: Specifying values larger than one makes only sense for some selection strategies. The argument is optional, default value is a vector consisting of ones. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
no.preselected |
Numerical vector specifying the minimum number of individuls that are preselected when selection for flanking marker (step "f") is carried out. Note: Only used for selection strategies including preselection for flanking markers. The argument is optional. If not specified, all individuals are preselected that show recombination at the maximum number of flanking markers. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
reg.chr |
Definition of chromosome regions, which are evaluated separately for
the recurrent parent genome. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
reg.begin |
Definition of chromosome regions, which are evaluated separately for
the recurrent parent genome. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
reg.end |
End of the region to be evaluated, in cM distance from the
telomere. If 0 is given, the entire chromosome is evaluated. The default
is | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
missing.allele |
Character string that specifies missing marker data.
The default is | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
result.file |
Required name of the file in which the simulation result is stored. Relative names are resolved against | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
success.factor |
Numeric success probability used after selection. The default is 1. If it is below 1, the number of successful selected individuals is drawn from a binomial distribution and at least one individual is retained. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
recode.infiles |
Logical. If | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
recode.file |
File name for recoded parental marker data when |
Value
On successful completion, invisibly returns a named list containing both the simulation definition and the simulation summaries. The parameter part includes simulation.name, simulation.run, repetitions, recurrent.parent, donor.parent, linkage.map, target.loci, flanking.loci, recipient.loci, gen.type, population.size, sel.strategy, no.selected, no.preselected, reg.chr, reg.begin, reg.end, and success.factor.
Depending on the requested analyses, result components include TargetAlleles, RPG.gw, numbered RPG.reg components, DonorSegments.gw, numbered DonorSegments.reg components, numbered LinkageDrag components, SM, and HT. Validation failures return NULL invisibly. The same list is saved in the explicitly supplied result.file.
See Also
mab.compare, mab.load.data,
mab.tabulate, mab.Input.files,
mab.Examples
Tabulate the results of a series of simulations
Description
Tabulates results from a series of simulations. The files to load are
supplied explicitly in result.files.
Usage
mab.tabulate(simulation.names.runs, par.summary = TRUE, ext.summary = FALSE,
RPG.gw.mean = FALSE, TargetAlleles = FALSE, RPG.gw.Q10 = FALSE,
RPG.reg.Q10 = FALSE, RPG.reg.mean = FALSE, DonorSegments.gw = FALSE,
DonorSegments.reg = FALSE, LinkageDrag = FALSE, MarkerAnalyses = FALSE,
digits = 1, percent = TRUE, result.files)
Arguments
simulation.names.runs |
Character vector containing alternating simulation names and simulation-run identifiers. It is interpreted as a two-column matrix, one row per simulation result to be tabulated. |
par.summary |
Logical. Include a summary of the main simulation parameters. |
ext.summary |
Logical. Include an extended summary of the simulation parameters. |
RPG.gw.mean |
Logical. Include mean genome-wide recurrent-parent genome (RPG) values. |
TargetAlleles |
Logical. Include the mean target-allele summary. |
RPG.gw.Q10 |
Logical. Include the Q10 value of the genome-wide RPG. |
RPG.reg.Q10 |
Logical. Include Q10 values of RPG in the defined chromosome regions. |
RPG.reg.mean |
Logical. Include mean RPG values in the defined chromosome regions. |
DonorSegments.gw |
Logical. Include the mean number of donor segments genome-wide. |
DonorSegments.reg |
Logical. Include the mean number of donor segments in the defined chromosome regions. |
LinkageDrag |
Logical. Include mean linkage-drag summaries for the target loci. |
MarkerAnalyses |
Logical. Include mean single-marker (SM) and high-throughput (HT) marker-analysis counts. |
digits |
Number of decimal places used when rounding numerical summaries. |
percent |
Logical. If |
result.files |
Character vector containing one explicit result-file name for each scenario in |
Value
A named list of the summary matrices requested by the logical arguments. Possible components include par.summary, ext.summary, TargetAlleles.mean, RPG.gw.mean, RPG.gw.Q10, numbered regional RPG.reg*.mean and RPG.reg*.Q10 matrices, DonorSegments.gw.mean, numbered DonorSegments.reg*.mean matrices, numbered LinkageDrag*.mean matrices, SM.mean, and HT.mean. If incompatible scenario definitions are detected, NULL is returned invisibly.
See Also
mab.simulate, mab.load.data,
mab.compare, mab.Examples
Optimize a population using the package selection routines
Description
Optimize a population using the package selection routines.
Usage
optimize.population(PopNames)
Arguments
PopNames |
Population name or population specification. |
Value
Returns the same invisible implementation-level list as population.optimize(). The function is a compatibility wrapper and the meaningful result is the delegated population operation.
Create a linkage map for phenotype or simulation use
Description
Create a linkage map for phenotype or simulation use.
Usage
ph.linkage.map.create(map, disperse=FALSE, file.disperse=NA,
disperse.factor=100)
Arguments
map |
Linkage map or map object. |
disperse |
Function argument used by this operation. |
file.disperse |
Function argument used by this operation. |
disperse.factor |
Function argument used by this operation. |
Value
When disperse = TRUE, returns the modified map object after positions have been dispersed. With the default disperse = FALSE, returns NULL; in both cases the main side effect is installing the map through the compiled map-definition routine.
Generate phenotypic values for a simulation population
Description
Generates phenotypic values for all loaded effects and all active individuals of a simulation population.
Usage
phenotype.population(PopName, hsq)
Arguments
PopName |
Character string naming one simulation population. |
hsq |
Numeric vector containing one heritability for each loaded effect,
in the order returned by |
Details
Phenotypic values are generated for all loaded effects in one operation.
Partial phenotypic evaluation is not supported. The length of hsq
must therefore equal the number of loaded effects.
The function first genotypes and evaluates the population. For effect
i, let V_{G,i} be the sample variance of the effect-specific
genetic values and let h_i^2 be the supplied heritability. The residual
variance is
V_{E,i} = V_{G,i}/h_i^2 - V_{G,i}.
For every individual an independent normal residual with variance
V_{E,i} is added to the effect-specific genetic value.
The phenotypic value array is parallel to the genetic value array. Elements 1 through the number of effects contain the effect-specific phenotypic values. Element 0 contains the weighted phenotypic selection index, calculated with the current effect weights in the same way as the genetic selection index.
Phenotypic values are population-wide derived data. Re-evaluating genetic values, explicitly setting genetic values, removing an evaluation, removing cached genotypes, changing the effect configuration or effect weights, or removing the linkage map invalidates stored phenotypic values.
Value
On successful phenotype generation, an invisible implementation-level list returned by the compiled routine; its retval component is the status code. Validation failures return NULL invisibly.
See Also
evaluate.population, get.population.pvalue,
population.sort, select.n.best, select.all.best
Perform the SelectionTools operation "plabsim"
Description
Perform the SelectionTools operation "plabsim".
Format
A list containing PLABSIM run-time state and parameters.
Value
A list containing the current PLABSIM run-time state and parameters.
Perform the SelectionTools operation "plabsim R date"
Description
Perform the SelectionTools operation "plabsim R date".
Format
A character string containing the PLABSIM R-code date.
Value
A character scalar containing the PLABSIM R-code date.
Perform the SelectionTools operation "plabsim R version"
Description
Perform the SelectionTools operation "plabsim R version".
Format
A character string containing the PLABSIM R-code version.
Value
A character scalar containing the PLABSIM R-code version.
Initialize the PLABSIM simulation state
Description
Initialize the PLABSIM simulation state.
Usage
plabsim.init()
Value
Returns the same invisible implementation-level list as rng.init(); the main effect is initializing the simulation random-number generator state.
Perform the SelectionTools operation "plabsim version"
Description
Perform the SelectionTools operation "plabsim version".
Usage
plabsim.version()
Value
Returns the same invisible implementation-level list as write.version.2(); the main effect is reporting package version information.
Create or process a plant object in the simulation framework
Description
Create or process a plant object in the simulation framework.
Usage
plant(genotype, NoInd=1)
Arguments
genotype |
Function argument used by this operation. |
NoInd |
Function argument used by this operation. |
Value
A data.frame with columns ind, chrom, hom, pos, and all, representing the requested genome from the supplied homologue genotypes.
Append a copy of a simulation population
Description
Appends a copy of a source population to the end of a destination population.
Usage
population.append(NameP1, NameP2)
Arguments
NameP1 |
Character string naming the recipient population. |
NameP2 |
Character string naming the source population, which is not modified. |
Details
Raw chromosome data are copied for every appended individual. If the destination is initially empty, complete genotype, genetic-value, and phenotypic-value categories are copied when present in the source. For a nonempty destination, a derived category is retained only when it is complete in both populations. Otherwise that category is cleared in the result.
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of appending a copy of one simulation population to another.
Concatenate simulation populations
Description
Moves all individuals of one population into another simulation population.
Usage
population.concat(NameP1, NameP2)
Arguments
NameP1 |
Character string naming the receiving population. |
NameP2 |
Character string naming the population to concatenate and remove. |
Details
Individuals are transferred as complete individual structures and
NameP2 is removed. As in the existing implementation, cached genotype
and genetic evaluation data are discarded before concatenation. Phenotypic
values are also discarded because they depend on the genetic evaluation.
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of concatenating simulation populations.
Copy a simulation population
Description
Copies all or a contiguous subset of a source population into a destination population.
Usage
population.copy(NameP1, NameP2, start=1, n=-1)
Arguments
NameP1 |
Character string naming the destination population. |
NameP2 |
Character string naming the source population. |
start |
One-based position of the first source individual to copy. |
n |
Number of individuals to copy. A negative value requests the remainder of the source population. |
Details
The source population is not modified. Raw chromosome data are copied. Complete cached genotype, genetic-value, and phenotypic-value categories are deep-copied when present in the source. Derived categories are maintained as all-or-none population data; allocation failure cannot leave only a subset of active destination individuals with a category.
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of copying a simulation population.
Divide a simulation population
Description
Moves the first requested individuals of a source population into a new population.
Usage
population.divide(NameP1, NameP2, NoI=-1)
Arguments
NameP1 |
Character string naming the destination population. |
NameP2 |
Character string naming the source population. |
NoI |
Number of individuals to move. |
Details
The operation moves complete individual structures. Therefore chromosome data, information fields, genotype data, genetic values, and phenotypic values remain attached to the same individuals. No value category is recalculated merely because the population is divided.
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of dividing a simulation population.
Check whether a simulation population exists
Description
Checks whether the requested population name occurs in the current simulation population list.
Usage
population.exist(PopName)
Arguments
PopName |
Name of the population. |
Value
An invisible logical scalar indicating whether PopName occurs in the current population list.
Individual remove a population or population data
Description
Individual remove a population or population data.
Usage
population.individual.remove(PopName, IndPos)
Arguments
PopName |
Name of the population. |
IndPos |
Function argument used by this operation. |
Value
For a valid individual position, the invisible implementation-level list returned by the final population operation used to remove the individual. If the position is outside the population, NULL. The main effect is modification of the population.
Return stored information for an individual
Description
Returns the character information stored for an individual in a simulation population.
Usage
population.info.get(PopName, ind)
Arguments
PopName |
Name of the population. |
ind |
Individual index or indices. |
Value
A character value containing the stored information for the requested individual when the individual index is valid; otherwise NULL.
Set stored information for an individual
Description
Sets the character information stored for an individual in a simulation population.
Usage
population.info.set(PopName, ind, info)
Arguments
PopName |
Name of the population. |
ind |
Individual index or indices. |
info |
Information value. |
Value
For a valid individual index, an invisible implementation-level list returned by the compiled setter. For an invalid index, NULL. The main effect is updating the stored individual information.
List a population or population data
Description
Lists the currently registered simulation populations and their active sizes.
Usage
population.list()
Details
The returned population registry is independent of SelectionTools marker-data-set names. The count associated with each name is the active number of individuals, not necessarily the allocated storage capacity.
Value
A data.frame with columns PopName and count, giving each current population name and its number of individuals.
Matrix load a population or population data
Description
Matrix load a population or population data.
Usage
population.matrix.load(file, backcross=FALSE, keep.IndID=TRUE)
Arguments
file |
File name or file connection. |
backcross |
Function argument used by this operation. |
keep.IndID |
Function argument used by this operation. |
Value
NULL. The function is called for its side effect of loading matrix-format population data and, when requested, restoring individual identifiers.
Matrix save a population or population data
Description
Matrix save a population or population data.
Usage
population.matrix.save(file, PopNames, missing.rm=TRUE, keep.IndID=TRUE)
Arguments
file |
File name or file connection. |
PopNames |
Population name or population specification. |
missing.rm |
Function argument used by this operation. |
keep.IndID |
Function argument used by this operation. |
Value
NULL, returned invisibly by the file-writing operation. The function is called for its side effect of writing matrix-format population data.
Swap the names of two simulation populations
Description
Swaps the registered names of two simulation populations.
Usage
population.name.swap(NameP1, NameP2)
Arguments
NameP1 |
Name of the first parental population. |
NameP2 |
Name of the second parental population. |
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of swapping simulation-population names.
Optimize storage for simulation populations
Description
Calls the compiled population-optimization routine for the requested simulation populations.
Usage
population.optimize(PopNames)
Arguments
PopNames |
Population name or population specification. |
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of optimizing the requested simulation populations.
Load a population from PLABSIM-style data
Description
Loads a simulation population from the PLABSIM-style file representation and restores stored individual information.
Usage
population.plabsim.load(file, PopName=NA)
Arguments
file |
File name or file connection. |
PopName |
Name of the population. |
Value
NULL. The function is called for its side effect of loading a population from the PLABSIM-style representation and restoring stored individual information.
Save a population in PLABSIM-style format
Description
Writes a simulation population and its stored individual information in the PLABSIM-style file representation.
Usage
population.plabsim.save(file, PopName)
Arguments
file |
File name or file connection. |
PopName |
Name of the population. |
Value
NULL. The function is called for its side effect of writing the population and individual information in the PLABSIM-style representation.
Remove a population or population data
Description
Removes one or more simulation populations.
Usage
population.remove(PopNames)
Arguments
PopNames |
Population name or population specification. |
Details
The named populations are deleted from the simulation population registry and their owned chromosome, information, genotype, and evaluation storage is released. Other simulation-wide definitions such as the genome, linkage map, and effects are separate and are not removed by this operation.
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of removing simulation populations.
Remove all a population or population data
Description
Removes all currently registered simulation populations.
Usage
population.remove.all()
Details
The function obtains the current population list and removes every listed population. This clears population-owned data but does not by itself redefine genome parameters, linkage maps, or effect maps.
Value
An invisible implementation-level list returned by the final information-message call after all populations have been removed. The meaningful result is the side effect of removing all populations.
Rename a simulation population
Description
Renames a registered simulation population.
Usage
population.rename(OldName, NewName)
Arguments
OldName |
Function argument used by this operation. |
NewName |
New population name. The name must contain at least two and fewer than 256 characters. |
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of renaming a simulation population.
Resize a population or population data
Description
Changes the allocated size of a simulation population.
Usage
population.resize(PopName, newSize)
Arguments
PopName |
Character string naming the population to resize. |
newSize |
Requested new population allocation. |
Details
A successful change of population allocation invalidates cached marker genotypes and evaluated genetic values because the active/storage relationship has changed. If no actual resize occurs, these derived categories are not cleared. If an attempted growth fails and the original allocation remains unchanged, the existing population and its derived data are preserved.
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of resizing a simulation population.
Sample a simulation population
Description
Creates a destination population by random sampling from a source population.
Usage
population.sample(NameP1, NameP2, size=-1, replace=FALSE)
Arguments
NameP1 |
Character string naming the sampled destination population. |
NameP2 |
Character string naming the source population. |
size |
Requested sample size. A negative value requests the source population size. |
replace |
Logical. If true, sampling is with replacement. |
Details
The source population is not changed. Raw chromosome data are copied for the sampled individuals. Complete cached genotype, genetic-value, and phenotypic-value categories are deep-copied when present in the source. Derived categories remain all-or-none for the destination population.
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of sampling a simulation population.
Return the size and allocated capacity of a simulation population
Description
Returns the active number of individuals and the number of individual slots allocated for a simulation population.
Usage
population.size.get(PopName)
Arguments
PopName |
Name of the population. |
Value
A one-row data.frame with integer columns NoInds (active number of individuals) and NoIndsAlloc (number of individuals for which memory is allocated). Returns NULL invisibly if the population is not available.
Sort a simulation population by genetic or phenotypic value
Description
Sorts the active individuals of a simulation population by a stored genetic or phenotypic value.
Usage
population.sort(PopName, decreasing=TRUE,
selection.criterion="G", effect=NULL)
Arguments
PopName |
Character string naming the population to sort. |
decreasing |
Logical. If true, larger values are placed first. |
selection.criterion |
Character string. |
effect |
Optional character string naming an effect. If omitted, element 0 of the selected value array is used. |
Details
With selection.criterion="G", the population is sorted by
GValue[0]. With selection.criterion="P", the corresponding
phenotypic value is used.
If effect is omitted, element 0 is the current weighted selection
index. If an effect is supplied, its effect-specific value is used.
All active individuals must have the requested value category. Sorting reorders complete individual structures, so chromosome data, genotype data, genetic values, phenotypic values, and information fields remain attached to the same individual.
Value
An invisible implementation-level list returned by the compiled sorter. Its retval component is the sort status; the main effect is reordering the requested population.
See Also
phenotype.population, select.n.best,
select.all.best
Swap simulation population names
Description
Compatibility wrapper for swapping the names of two simulation populations.
Usage
population.swap.name(NameP1, NameP2)
Arguments
NameP1 |
Name of the first parental population. |
NameP2 |
Name of the second parental population. |
Value
Returns the same invisible implementation-level list as population.name.swap(). The function is a compatibility wrapper and the meaningful result is the delegated population operation.
Transfer a population or population data
Description
Transfers the individuals of one simulation population into another, either retaining or deleting the source population.
Usage
population.transfer(NameP1, NameP2, population.NameP2.delete=TRUE)
Arguments
NameP1 |
Character string naming the recipient population. |
NameP2 |
Character string naming the source population. |
population.NameP2.delete |
Logical. If true, remove the source after transfer by concatenation; if false, retain the source and append a copy. |
Details
If population.NameP2.delete=TRUE, the operation uses population concatenation: the individuals of NameP2 are added to NameP1 and NameP2 is removed. If false, population append is used: the same individuals are copied to NameP1 while NameP2 remains available. Derived genotype and genetic-value state follows the consistency rules of the underlying append or concatenate operation.
Value
The invisible implementation-level list returned by population.concat() or population.append(), depending on population.NameP2.delete. The main effect is transfer of population data.
remove.all.populations()
Description
Removes all currently registered simulation populations.
Usage
remove.all.populations()
Details
The current population list is collected and all listed populations are removed. Genome parameters, linkage-map definitions, effect definitions, and marker-data sets are separate objects and are not removed with the simulation populations.
Value
Returns the same invisible implementation-level list as population.remove.all(); the main effect is removal of all populations.
remove.effmaps()
Description
Deletes all loaded effects
Usage
remove.effmaps()
Value
Returns the same invisible implementation-level list as effmap.remove.all(); the main effect is removal of all effect maps.
Remove genetic and phenotypic evaluation values
Description
Clears stored genetic and phenotypic values for one or more simulation populations.
Usage
remove.evaluate.population(PopNames)
Arguments
PopNames |
Population name or population specification. |
Details
Population membership, chromosome segments, and cached marker genotypes are retained. Both genetic values and phenotypic values are discarded. Genetic values must be evaluated or set again before genetic selection, and phenotypes must be generated again before phenotypic selection.
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of removing stored evaluation values.
Remove cached genotype information from a population
Description
Clears cached marker-genotype data for one or more simulation populations.
Usage
remove.genotype.population(PopNames)
Arguments
PopNames |
Population name or population specification. |
Details
The raw chromosome-segment representation is retained. Existing genetic values are retained. Stored phenotypic values are discarded, because phenotypic values are tied to the current genetic/genotype state. Genotype-dependent operations must regenerate marker genotypes from the current linkage map and chromosome data.
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of removing cached genotype information.
Remove the linkage map from the simulation
Description
Removes the current linkage map and map-dependent cached simulation data.
Usage
remove.map()
Details
Removing the linkage map clears cached marker genotypes, genetic values, phenotypic values, and loaded effect definitions because these data depend on the current map/effect configuration. Population chromosome-segment data are retained according to the simulation backend.
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of removing the current linkage map.
remove.population()
Description
Removes one or more simulation populations and all data owned by them.
Usage
remove.population(PopNames)
Arguments
PopNames |
Names of the populations to be deleted |
Details
Removing a population frees its chromosome representation and associated per-population storage, including cached genotype and evaluation data. Multiple population names can be supplied.
Value
Returns the same invisible implementation-level list as population.remove(). The function is a compatibility wrapper and the meaningful result is the delegated population operation.
rename.population()
Description
Changes the registered name of a simulation population without changing its individuals.
Usage
rename.population(OldName, NewName)
Arguments
OldName |
Name of the population to be renamed |
NewName |
New name of the population |
Details
The population object and its chromosome, genotype, evaluation, and information data remain the same; only the population identifier changes.
Value
Returns the same invisible implementation-level list as population.rename(). The function is a compatibility wrapper and the meaningful result is the delegated population operation.
reset.all()
Description
"Resets" the program into the initial state.
Usage
reset.all()
Value
Returns the same invisible implementation-level list as reset.mdp() after resetting the other simulation state. The main effect is resetting all package simulation state handled by the function.
Note
All informations are lost, besides the genome parameter
(genome.parameter.set).
Reset marker-assisted donor-parent information
Description
Reset marker-assisted donor-parent information.
Usage
reset.mdp()
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of resetting marker-assisted donor-parent information.
Resize a population
Description
Changes the allocated size of a simulation population.
Usage
resize.population(PopName, newSize)
Arguments
PopName |
Character string naming the population to resize. |
newSize |
Requested new population allocation. |
Details
A successful change of allocation invalidates cached marker genotypes and evaluated genetic values because these derived categories must remain complete and consistent for the active population. If no actual resize occurs, the cached derived data are left unchanged. If an attempted growth fails and the original allocation remains unchanged, the original population state is preserved.
Value
Returns the same invisible implementation-level list as population.resize(). The function is a compatibility wrapper and the meaningful result is the delegated population operation.
Return or display resource information used by SelectionTools
Description
Return or display resource information used by SelectionTools.
Usage
resources(expr)
Arguments
expr |
Function argument used by this operation. |
Value
Invisibly returns the value produced by evaluating expr; its class and structure are therefore exactly those returned by expr. Resource/timing information is reported as a side effect.
return.population()
Description
Converts one or more simulation populations to the marker-incidence data-frame representation used by the genetic-distance routines.
Usage
return.population(PopNames, missing.rm=TRUE)
Arguments
PopNames |
Character string containing one or more simulation population names separated by blanks. |
missing.rm |
Logical flag controlling removal or retention of missing marker information in the returned representation. |
Details
The requested simulation populations are evaluated at the current marker loci and returned in an NTSys-style matrix representation suitable for the corresponding SelectionTools data utilities. Multiple population names can be supplied in one space-separated string.
The returned object is a representation of population marker data; it does not remove or modify the source populations. The missing.rm argument controls handling of missing marker information in the conversion.
Value
A data.frame in the marker-incidence representation used by the genetic-distance routines. Returns NULL invisibly if the compiled population conversion cannot produce a result.
Choose the simulation random-number generator state
Description
Choose the simulation random-number generator state.
Usage
rng.choose(rngName="")
Arguments
rngName |
Function argument used by this operation. |
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of selecting the requested random-number generator.
Info the simulation random-number generator state
Description
Info the simulation random-number generator state.
Usage
rng.info()
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of requesting random-number-generator information.
Init the simulation random-number generator state
Description
Init the simulation random-number generator state.
Usage
rng.init()
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of initializing the random-number generator.
List the simulation random-number generator state
Description
List the available simulation random-number generators and indicate the current choice. Output follows the package information-level setting and can therefore be suppressed through that setting.
Usage
rng.list()
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of requesting the available random-number generators.
sample.population()
Description
Creates a new population by sampling individuals from an existing simulation population.
Usage
sample.population(NameP1, NameP2, NoI=-1, rep=1)
Arguments
NameP1 |
Name of the sample population |
NameP2 |
Name of the initial population |
NoI |
non-negative integer giving the number of individuals to choose |
rep |
Should sampling be with replacement? |
Details
Sampling does not modify the source population. The requested sample is copied into the destination. Sampling can be with or without replacement according to rep. Raw chromosome data are always the authoritative copied representation; cached genotype and genetic-value categories are transferred only when complete in the source and successfully constructed for the entire destination.
Value
Returns the same invisible implementation-level list as population.sample(). The function is a compatibility wrapper and the meaningful result is the delegated population operation.
Save a linkage map
Description
Save a linkage map.
Usage
save.linkage.map(file, dta.map=get.map())
Arguments
file |
File name or file connection. |
dta.map |
Map-related parameter. |
Value
NULL, returned by the delegated linkage.map.save() file-writing operation. The function is a compatibility wrapper.
Calculate a standard-deviation quantity used by the simulation code
Description
Calculate a standard-deviation quantity used by the simulation code.
Usage
sdev(x, na.rm = FALSE)
Arguments
x |
Input object. |
na.rm |
Function argument used by this operation. |
Value
A numeric standard-deviation result. For a matrix or data frame, a numeric vector is returned with one standard deviation per column; for a vector or other object coerced to a vector, a numeric scalar is returned.
Define a selection target
Description
Define a selection target.
Usage
sel.target.define(criteria.df)
Arguments
criteria.df |
Function argument used by this operation. |
Value
A character vector containing one or two selection-target codes selected from "H1", "MIX", "H2", and "All". Returns NULL when the criterion cannot be translated into a supported target.
Select all individuals in the best value classes
Description
Selects individuals belonging to the best distinct genetic or phenotypic value classes of a simulation population.
Usage
select.all.best(newPop, oldPop, effect=NULL, x=1,
decreasing=TRUE, selection.criterion="G")
Arguments
newPop |
Character string naming the selected population. |
oldPop |
Character string naming the source population. |
effect |
Optional character string naming the effect used for selection. |
x |
Number of distinct value classes to retain. |
decreasing |
Logical. If true, larger values are preferred. |
selection.criterion |
Character string. |
Details
With selection.criterion="G", the source population is genotyped and
evaluated before sorting. Without a named effect, the weighted genetic index
is used; with a named effect, selection is based on that effect.
Phenotypic selection uses previously stored phenotypic values and does not
re-evaluate the population. Without a named effect, the weighted phenotypic
index is used; with a named effect, its effect-specific phenotype is used.
The argument x counts distinct values, so ties are retained together.
Continuous phenotypic values will ordinarily produce fewer ties than many
genetic scoring schemes.
Value
On successful selection, invisibly returns a one-row data.frame with columns n, minscore, and maxscore, giving the number selected and the score limits of the selected class(es). Returns NULL for an empty/unusable selection, or the invisible status list from population.sort() if sorting fails.
See Also
phenotype.population, population.sort,
select.n.best
Determine the best value classes in a sorted population
Description
Determines how many individuals belong to the requested number of best value classes in an already sorted simulation population.
Usage
select.all.best.intern(NamePop, n=1, selection.criterion="G", effect=NULL)
Arguments
NamePop |
Character string naming the already sorted population. |
n |
Number of distinct value classes to retain. |
selection.criterion |
Character string. |
effect |
Optional character string naming the effect whose genetic or phenotypic value is used. If omitted, element 0 of the selected value array is used. |
Details
The function does not sort the population. It is used internally after
population.sort. The first n distinct values in the current
population order are treated as the best classes, and ties within a class
are retained together.
With selection.criterion="G", genetic values are inspected. With
selection.criterion="P", previously generated phenotypic values are
inspected. If effect is omitted, element 0 of the corresponding value
array is used; otherwise the named effect-specific value is used.
Value
On success, a one-row data.frame with columns n, minscore, and maxscore, giving the number of individuals in the selected class(es) and the corresponding score limits. Returns NULL invisibly if the compiled selection step fails.
See Also
select.all.best, population.sort
Select specific allele combinations from populations
Description
This function is used to select individuals from a population that fulfill certain selection criteria.
Usage
select.genotypes(newPop, oldPop, criteria)
Arguments
newPop |
a character string specifying the name of the new population which is created from the selected genotype(s). |
oldPop |
a character string specifying the name of the population that selection is applied to. |
criteria |
a data.frame containing the locinames and the selection criteria for selection. Every row is one criterion. Data.frame must contain three columns: locusname, allele 1, allele 2. |
Details
Individuals are selected based on their genotypes at the loci listed in 'criteria'. Genotype matching is performed using internally created effects.
If multiple rows are provided for the same 'locus', an individual matches that locus if it satisfies **any** of the rows for that locus (OR within locus). Individuals must satisfy the requirements for **all** loci present in 'criteria' (AND across loci).
Selected individuals are removed from 'oldPop' and added to 'newPop'.
Value
On successful completion, returns the invisible implementation-level list produced by the final temporary-population removal. If the source population is empty, NULL is returned invisibly. The meaningful result is the side effect of constructing newPop according to the requested allele criteria.
See Also
[select.all.best()], [select.n.best()], [define.effects()]
Examples
reset.all()
map <- data.frame(1,0.0,"form","trait1")
define.genome(map)
init.population("P1",homozygote(1,100))
init.population("P2",homozygote(2,100))
append.population("ADM","P1")
append.population("ADM","P2")
cross("SYN1","ADM","ADM",1000)
evaluate.genotype("SYN1","form")
evaluate.allele.freq("SYN1","form")
# Select genotypes (1|1) and (1|2)
sel.crit <- data.frame(
locus = c ("form" , "form"),
allele1 = c (1 , 1 ),
allele2 = c( 1 , 2 )
)
select.genotypes("selected","SYN1",
sel.crit)
evaluate.genotype("selected","form")
Select a fixed number of best individuals
Description
Selects a fixed number of individuals from a simulation population according to genetic or phenotypic values.
Usage
select.n.best(newPop, oldPop, effect=NULL, n=1,
decreasing=TRUE, selection.criterion="G")
Arguments
newPop |
Character string naming the selected population. |
oldPop |
Character string naming the source population. |
effect |
Optional character string naming the effect used for selection. |
n |
Number of individuals requested. |
decreasing |
Logical. If true, larger values are preferred. |
selection.criterion |
Character string. |
Details
With selection.criterion="G", the source population is genotyped and
evaluated before selection. With effect=NULL, selection is based on
the weighted genetic index.
With a named effect, selection is based on that effect.
With selection.criterion="P", stored phenotypic values are used and
the population is not re-evaluated, because genetic re-evaluation invalidates
phenotypic values. If effect=NULL, the weighted phenotypic index is
used; with a named effect, its effect-specific phenotype is used.
Value
On successful selection, returns the invisible implementation-level list produced by population.divide(); if sorting fails the sort-status list is returned invisibly, and an empty population returns NULL invisibly. The main effect is creation of newPop containing the selected individuals.
See Also
phenotype.population, population.sort,
select.all.best
Select the best genome segments according to the requested criterion
Description
Select the best genome segments according to the requested criterion.
Usage
select.n.best.segments(newPop, oldPop, n=1, allele = 1, decreasing = FALSE)
Arguments
newPop |
Function argument used by this operation. |
oldPop |
Function argument used by this operation. |
n |
Number of items or repetitions. |
allele |
Allele specification. |
decreasing |
Function argument used by this operation. |
Value
On successful selection, returns the invisible implementation-level list produced by population.divide(); if the requested number cannot be selected, NULL is returned invisibly. The main effect is creation of newPop from individuals ranked by the number of genome segments.
Set the initial number of loci used by the simulation
Description
Set the initial number of loci used by the simulation.
Usage
set.NoLociInit(NoLociInit)
Arguments
NoLociInit |
Initial number of loci. |
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of setting the initial number of loci.
Set crossover-frequency information
Description
Set crossover-frequency information.
Usage
set.co.freq(cofreq=1)
Arguments
cofreq |
Function argument used by this operation. |
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of setting crossover-frequency information.
set.eff.weight()
Description
Assigns a weight to a defined effect
Usage
set.eff.weight(fname, weight)
Arguments
fname |
Name of the effect |
weight |
Weight of the effect |
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of setting the requested effect weight.
Note
Effect weights are used when evaluate.genotype is
carried out without specifing one effect. In this case all loaded
effects are used to determin tne genotypic value, and weights are
applied.
If evaluate genotype is used with specifiying an effect, the
weights are not used.
See define.effects for for an example
set.genome.par()
Description
Defines simulation chromosome number, ploidy, and chromosome lengths.
Usage
set.genome.par(no.chrom, no.hom, chrom.len)
Arguments
no.chrom |
Positive number of simulation chromosomes. |
no.hom |
Number of homologues. Diploid population operations use two. |
chrom.len |
Numeric vector of chromosome lengths with one value per chromosome. |
Details
The genome definition controls the chromosome structure used by simulation-population operations. no.chrom gives the number of chromosomes, no.hom the number of homologues, and chrom.len the chromosome lengths in the simulation map unit.
Population-generation operations including crossing, single-seed descent, doubled haploids, and the SelectionTools marker-to-simulation bridge require a diploid genome with two homologues. The chromosome-length vector must correspond to the declared chromosome count.
This function and genome.parameter.set modify the same simulation genome state.
Value
Returns the same implementation-level list as genome.parameter.set(); the main effect is setting the simulation genome parameters.
set.info.level()
Description
The function set.info.level() sets the information level. The smaller the level number is the less program information is printed on the screen.
| level | information printed |
| -2 | error messages |
| -1 | warning messages |
| 0 | default messages |
| 1 | verbouse mode |
Usage
set.info.level(level)
Arguments
level |
information level |
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of setting the simulation information level.
Note
If you want to read only error messages then you have to set your information level to -2!
Set marker-assisted donor-parent information
Description
Set marker-assisted donor-parent information.
Usage
set.mdp(MDP=0)
Arguments
MDP |
Function argument used by this operation. |
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of setting marker-assisted donor-parent information.
Set genetic selection values for a simulation population
Description
Sets element 0 of the stored genetic value array for every active individual of a simulation population.
Usage
set.population.gvalue(name, gvalue)
Arguments
name |
Character string naming the population. |
gvalue |
Finite numeric vector containing exactly one value for every active individual. |
Details
The length of gvalue must exactly equal the active population size.
The operation is population-wide: values are set for all active individuals
or the operation is unsuccessful. Partial assignment and missing-value
padding are not used.
When effects are loaded, the R interface first genotypes and evaluates the population so that the complete genetic value arrays exist for all effects. The supplied values then replace element 0 only; the effect-specific genetic values remain available. When no effects are loaded, the complete genetic value array consists only of element 0.
A successful assignment invalidates all stored phenotypic values for the population.
Value
On successful assignment, an invisible implementation-level list returned by the compiled setter, including its retval status component. Invalid population, length, or non-finite input returns NULL invisibly.
See Also
get.population.gvalue, evaluate.population,
phenotype.population
Set population information used by the simulation routines
Description
Set population information used by the simulation routines.
Usage
set.population.info(PopName, ind, info)
Arguments
PopName |
Name of the population. |
ind |
Individual index or indices. |
info |
Information value. |
Value
Returns the same invisible implementation-level list or NULL as population.info.set(), depending on whether the individual index is valid.
Create or evaluate a single cross
Description
Create or evaluate a single cross.
Usage
single.cross(PgName, PopName, classSize=1)
Arguments
PgName |
Function argument used by this operation. |
PopName |
Name of the population. |
classSize |
Function argument used by this operation. |
Value
An invisible implementation-level list returned by the final population-removal operation. The meaningful result is the side effect of creating the requested single-cross population.
Start or stop the simulation timing helper
Description
Start or stop the simulation timing helper.
Usage
sm.start.timer(depth=1)
Arguments
depth |
Function argument used by this operation. |
Value
An invisible implementation-level list returned by the SelectionTools timing helper. The function is called for the side effect of starting timing.
Start or stop the simulation timing helper
Description
Start or stop the simulation timing helper.
Usage
sm.stop.timer(depth=1, info.level=0)
Arguments
depth |
Function argument used by this operation. |
info.level |
Function argument used by this operation. |
Value
An invisible implementation-level list returned by the SelectionTools information helper after reporting elapsed time. The function is called for the timing/reporting side effect.
Split a data object according to the package-specific convention
Description
Split a data object according to the package-specific convention.
Usage
splitdt(m.a)
Arguments
m.a |
Marker or matrix input. |
Value
A character matrix with two columns and one row for each input string. The first column contains the part before the first dot (or underscore if no dot is present), and the second column contains the remaining part; if no separator is present the second field is empty.
ssd.mating()
Description
Generates single-seed-descent progeny by repeated selfing of each individual in a diploid parental population.
Usage
ssd.mating(NamePg, NameP, NoPg=1, maxcycles=1)
Arguments
NamePg |
Character string naming the progeny population to create. |
NameP |
Character string naming the parental population. |
NoPg |
Positive number of independently generated descendants per parental individual. |
maxcycles |
Positive maximum number of successive selfing generations for each descendant. |
Details
For every active individual in NameP, NoPg descendant lines are generated. The total progeny size is therefore the parental population size multiplied by NoPg. Descendants belonging to one parent occupy consecutive positions in the new population.
Each descendant is repeatedly selfed for at most maxcycles generations. Selfing stops earlier when the simulated chromosomes have become completely homozygous. Large values of maxcycles can therefore be used to approach fully homozygous recombinant inbred lines without requiring unnecessary generations after homozygosity is reached.
The routine requires a diploid simulation genome. The destination is rebuilt from chromosome data and does not retain stale cached genotypes or evaluated values. Allocation or meiosis failure removes the incomplete destination population.
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of creating progeny using the SSD mating routine.
Plot a Linkage Disequilibrium Heat Map
Description
Plots a triangular heat map of pairwise linkage disequilibrium (LD) values using only the standard R graphics system. Marker cells are equally spaced in the LD matrix. If marker positions are supplied, their genetic or physical spacing is displayed separately on a map line connected to the matrix.
Usage
st.LDheatmap(LD, map = NULL,
distances = "genetical",
LDmeasure = "r",
title = "",
col = grDevices::heat.colors(20),
zlim = c(0, 1),
add.map = TRUE,
add.key = TRUE,
depth = Inf,
max.dist = Inf,
write.ld = NULL,
write.snp.id = FALSE,
text = FALSE,
digits = 2,
geneMapLocation = 0.15,
geneMapLabelX = NULL,
geneMapLabelY = NULL,
cex.title = 0.9,
cex.map = 0.7,
cex.key = 0.6,
cex.ld = NULL,
cex.snp = NULL)
Arguments
LD |
A numeric square matrix containing pairwise LD values. The marker order
is given by the row and column order of the matrix. Values from the upper
triangle are plotted. If an upper-triangle entry is |
map |
A numeric vector containing one map position for each marker, in the same
order as the rows and columns of |
distances |
Character string specifying the units represented by |
LDmeasure |
Character string used to label the color key. Values |
title |
A single character string giving the plot title. An empty string omits the title. |
col |
A vector of at least two valid R colors. The first supplied color is used
for the highest LD values and the last supplied color for the lowest LD
values. Thus |
zlim |
Numeric vector of length two giving the lower and upper limits of the LD
color scale. The default is |
add.map |
Logical. If |
add.key |
Logical. If |
depth |
A non-negative integer or |
max.dist |
A non-negative number or |
write.ld |
|
write.snp.id |
Logical. If |
text |
Logical convenience option. If |
digits |
Non-negative integer giving the number of decimal places used when
|
geneMapLocation |
Non-negative numeric value controlling the perpendicular distance of the map line from the matrix diagonal. Larger values move the map line farther from the matrix. |
geneMapLabelX |
Optional finite numeric x-coordinate for the text reporting the total map
length. If |
geneMapLabelY |
Optional finite numeric y-coordinate for the text reporting the total map
length. If |
cex.title |
Positive finite numeric character expansion factor for the plot title. |
cex.map |
Positive finite numeric character expansion factor for the map-length label. |
cex.key |
Positive finite numeric character expansion factor for the color-key title and tick labels. |
cex.ld |
|
cex.snp |
|
Details
st.LDheatmap is a plotting function; it does not calculate LD. A
precomputed pairwise LD matrix is supplied in LD. In a typical
SelectionTools workflow the matrix and map can be obtained with
st.LDplot.ld and st.LDplot.map, respectively.
The heat-map cells represent marker order rather than map distance. The marker
centers are therefore equally spaced from zero to one along both matrix axes.
When add.map = TRUE, the actual values in map are rescaled to a
separate line parallel to the matrix diagonal, and each matrix position is
connected to its corresponding map position. This keeps the LD cells equally
sized while still displaying uneven marker spacing. Markers with identical map
positions are supported and are connected to the same position on the map
line.
Only the upper triangle of LD is displayed, and the diagonal is omitted.
An upper-triangle value is authoritative when it is present. If it is missing
and the corresponding lower-triangle value is non-missing, that lower-triangle
value is copied to the displayed upper triangle. Missing LD values remain blank.
The lower triangle is otherwise ignored.
The restrictions specified by depth and max.dist are applied
simultaneously. A cell is displayed only if it satisfies both restrictions.
The restrictions affect only which LD cells are drawn; they do not change the
marker order or the map line.
The color scale covers zlim. The order of col is from high LD to
low LD. This convention makes heat.colors() produce the customary red
for high LD and yellow for low LD. Missing or filtered cells are not painted.
If write.ld is supplied, it is evaluated once for each displayed LD
value. This permits selective labels without changing the plotted matrix. If
cex.ld = NULL, the function estimates a label size intended to fit into
one heat-map cell. Automatic sizing is necessarily approximate on very dense
plots; an explicit cex.ld can be supplied when exact control is needed.
The same consideration applies to marker identifiers and cex.snp.
The function uses standard R graphics and starts one new plot in the current
graphics layout. It does not modify global par settings. Consequently,
standard layouts such as par(mfrow = c(2, 2)) can be used to place
several LD heat maps on one graphics device.
Value
An invisible named list with components:
LDmatrix |
The upper-triangular LD matrix actually displayed after the plotting restrictions are applied. |
map |
The supplied map vector, or |
snp.position |
Plotting coordinates of markers in the LD matrix. |
map.x, map.y |
Plotting coordinates for the map line, or |
distances |
The map-distance type used internally. |
depth |
The marker-depth restriction. |
max.dist |
The map-distance restriction. |
zlim |
The color-scale limits. |
col |
The color vector used for the heat map. |
See Also
st.calc.ld, st.LDplot.ld, st.LDplot.map
Examples
LD <- matrix(c(
1.00, 0.85, 0.35, 0.10, 0.05,
0.85, 1.00, 0.70, 0.25, 0.10,
0.35, 0.70, 1.00, 0.80, 0.30,
0.10, 0.25, 0.80, 1.00, 0.90,
0.05, 0.10, 0.30, 0.90, 1.00
), nrow = 5, byrow = TRUE)
rownames(LD) <- colnames(LD) <- paste0("m", 1:5)
map <- c(0, 5, 17, 18, 35)
st.LDheatmap(LD, map,
title = "Pairwise LD",
col = grDevices::heat.colors(20))
st.LDheatmap(LD, map,
depth = 2,
write.ld = function(x)
if (x >= 0.7) sprintf("%.2f", x) else "",
write.snp.id = TRUE)
st.LDheatmap(LD, map,
max.dist = 10,
add.key = FALSE)
Prepare linkage-disequilibrium information for plotting
Description
Prepare linkage-disequilibrium information for plotting.
Usage
st.LDplot.ld(ld, chrom)
Arguments
ld |
Function argument used by this operation. |
chrom |
Chromosome identifier. |
Value
A numeric matrix containing the linkage-disequilibrium values arranged for plotting.
Prepare linkage-disequilibrium information for plotting
Description
Prepare linkage-disequilibrium information for plotting.
Usage
st.LDplot.map(chrom, data.set)
Arguments
chrom |
Chromosome identifier. |
data.set |
Name of the SelectionTools data set. |
Value
A numeric vector containing map positions for the requested chromosome.
Convert Compact STvcf Data to a VCF-Like Data Frame
Description
Expands a compact STvcf list to a wide VCF-like R data frame with one marker per row and one genotype column per individual.
Usage
st.STvcf.to.dataframe(STvcf)
Arguments
STvcf |
A compact STvcf list containing marker metadata, individual identifiers, chromosome mapping, and two allele matrices. |
Details
The function validates the STvcf object before expansion. The list must
contain exactly the nine documented STvcf components, each exactly once;
extra, missing, duplicated, unnamed, or unknown components are rejected.
Component types and dimensions must agree, chromosome codes must be
consecutive and used, map
positions must be finite and nonnegative, allele indices must agree with
REF/ALT, and marker and individual identifiers must obey the same restrictions
as st.load.vcf.data().
Marker and individual identifiers must be unique, shorter than 256 bytes, and contain only ASCII letters and digits. Dots and all other punctuation are rejected. These checks are intentionally the same on conversion in both directions.
REF may not be missing, ".", empty, or comma-separated. ALT may be
"." or a comma-separated list. ALT entries must be non-empty, unique,
and different from REF. Invalid or duplicated allele labels are rejected.
The returned data frame contains CHROM, POS, ID,
REF, ALT, and FORMAT, followed by one sample column per
entry in STvcf$individual. FORMAT is written as "GT" for every
marker. Including FORMAT is deliberate: it makes the sample boundary
unambiguous even when an individual identifier is itself an alphanumeric name
such as POS, INFO, CM, or FORMAT.
CHROM is written using CHROM.levels. POS is written as the current
STvcf double-precision genetic position in cM without intentional rounding.
Genotypes use VCF allele indices, with 0 for REF and 1, 2, ... for ALT
alleles. Allele order is retained. Every non-missing genotype is written with
"/"; "|" is never written because STvcf does not retain the
original separator and must not imply that phasing was inferred. If either
stored allele is missing, the genotype is written as "./.".
The two allele matrices must use the same storage type. Raw matrices use raw
value 255 for missing data; integer matrices use NA_integer_ for
missing data. STvcf objects accepted by SelectionTools may define at most 253
ALT alleles per marker, so valid non-missing VCF allele indices range from 0
through 253. Allele indices are checked against the corresponding REF/ALT
definition.
STvcf does not retain QUAL, FILTER, INFO, non-GT FORMAT fields, phase-set
metadata, or a physical POS field that was superseded by CM. These fields
cannot be reconstructed. For an STvcf object produced by
st.dataframe.to.STvcf(), conversion to a data frame and back preserves
the STvcf information under the documented normalizations: GT separators are
written as slash and missing genotypes as "./.".
Value
A VCF-like data.frame with fixed columns CHROM, POS, ID, REF, ALT, and FORMAT, followed by one genotype column for each individual.
Calculate linkage disequilibrium from marker data
Description
Calculate linkage disequilibrium from marker data.
Usage
st.calc.ld(ld.measure = "r2", auxfiles = TRUE, data.set = "default")
Arguments
ld.measure |
Function argument used by this operation. |
auxfiles |
Logical. If |
data.set |
Name of the SelectionTools data set. |
Value
Invisibly returns a data.frame with columns Chrom, Locus1, Locus2, Name1, Name2, and LD, containing pairwise linkage-disequilibrium results; otherwise NULL.
Calculate linkage disequilibrium from marker data
Description
Calculate linkage disequilibrium from marker data.
Usage
st.calc.ld.2(ld.measure = "r2", auxfiles = TRUE, keep.ldfile = FALSE,
data.set = "default", ld.filename = NULL)
Arguments
ld.measure |
Function argument used by this operation. |
auxfiles |
Logical. Controls generation of the LD result file. With |
keep.ldfile |
Logical. If |
data.set |
Name of the SelectionTools data set. |
ld.filename |
Character string naming the persistent LD output file. Required only when |
Value
Invisibly returns a data.frame with columns Chrom, Locus1, Locus2, Name1, Name2, and LD when the generated LD file is read back. If keep.ldfile = TRUE, the R return value is NULL; when auxfiles = TRUE, the LD file named by ld.filename is retained.
Calculate the package-specific Q statistic or quantity
Description
Calculate the package-specific Q statistic or quantity.
Usage
st.calc.q(pop.type = "DH", t = 0, auxfiles = TRUE, data.set ="default")
Arguments
pop.type |
Population type. |
t |
Generation or method-specific time parameter. |
auxfiles |
Logical. If |
data.set |
Name of the SelectionTools data set. |
Value
Invisibly returns a data.frame with columns Chrom, Locus1, Locus2, and q, containing the calculated pairwise q values; otherwise NULL.
Calculate recombination fractions
Description
Calculate recombination fractions.
Usage
st.calc.rf(map.function = "Haldane", auxfiles = TRUE, data.set = "default")
Arguments
map.function |
Map function to use. |
auxfiles |
Logical. If |
data.set |
Name of the SelectionTools data set. |
Value
Invisibly returns a data.frame with columns Chrom, Locus1, Locus2, and RecFreq, containing pairwise recombination fractions; otherwise NULL.
Return chromosome statistics for a data set
Description
Return chromosome statistics for a data set.
Usage
st.chrom.stats(filename = "chrom.stats", auxfiles = TRUE, data.set ="default")
Arguments
filename |
Character string retained for compatibility and used as part of the name of the internal session-temporary result file. |
auxfiles |
Logical. If |
data.set |
Name of the SelectionTools data set. |
Value
Invisibly returns a data.frame with columns Chrom, NLoci, and Length, giving chromosome identifier, number of loci, and chromosome length; otherwise NULL if the calculation fails.
st.copy.marker.data()
Description
Creates an independent copy of the primary marker data, marker statistics, linkage-map representation, and phenotype vector of another SelectionTools data set.
Usage
st.copy.marker.data(target.data.set, source.data.set)
Arguments
target.data.set |
Character string naming the data set that will receive the copy. Existing contents of this target are replaced. |
source.data.set |
Character string naming the marker data set to copy. |
Details
The target receives its own marker-name and individual-name arrays, allele matrices, recomputed marker statistics, map storage, map access pointers, marker-to-map indices, and phenotype vector when present. Map storage order and map access order are reconstructed so that the copy does not depend on pointers into the source data set.
The copy is intended as a copy of the marker-data state, not of every fitted model or temporary calculation attached to the source. Model matrices, fitted effects, p-values, cross-validation state, and other derived analysis objects are not implied by this operation.
The source data set is not modified. Existing contents of the target data set are replaced.
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of copying marker data from the source data set to the target data set.
Construct a path in the SelectionTools data directory
Description
Construct a path in the SelectionTools data directory.
Usage
st.datadir(filename)
Arguments
filename |
File name. |
Value
A character scalar containing the constructed path in the SelectionTools data directory.
Convert a VCF-Like Data Frame to Compact STvcf Format
Description
Converts a wide diploid VCF-like R data frame to the compact STvcf list
representation used by st.load.vcf.data().
Usage
st.dataframe.to.STvcf(data)
Arguments
data |
A data frame containing VCF-style fixed fields and sample genotypes. The required layout and conversion conventions are described below. |
Details
The converter is deliberately strict. It is intended to create only STvcf objects that can subsequently be accepted by the SelectionTools marker-data and simulation-transfer path. Invalid marker or individual names, malformed genotypes, inconsistent REF/ALT definitions, invalid map positions, and unsupported allele indices are rejected during conversion rather than being left for a later consumer to discover.
The first five data-frame columns are fixed by position and must be, in this
order, CHROM or #CHROM, POS, ID, REF,
and ALT. This positional convention is intentional. Sample names are
allowed to be ordinary alphanumeric strings such as POS, INFO,
or FORMAT; therefore fixed fields cannot safely be identified only by
searching all column names.
Data-frame column names must be non-missing and non-empty. Duplicate names are permitted only when they arise because a sample has the same name as a fixed or metadata field. Sample names themselves must be unique.
Marker IDs are mandatory and unique. Marker and individual identifiers may
contain only ASCII letters A-Z and a-z and digits
0-9. Dots, underscores, hyphens, whitespace, punctuation, and
non-ASCII characters are rejected. Names must contain fewer than 256 bytes.
The same restrictions are checked again by st.load.vcf.data() and by
the reverse marker-data conversion. Names are never silently modified.
Input chromosome labels may be numeric or character. They are converted to
consecutive running integer codes 1, 2, ..., in order of first appearance.
The original input labels are retained in CHROM.levels. One
information message at level 0 reports the mapping. Chromosome labels are not
marker or individual names and are not subject to the alphanumeric name
restriction.
STvcf positions are genetic positions in centimorgan. If a valid metadata
column named CM is supplied, it is used for the genetic position and
the input POS values are not used for mapping. Otherwise
POS is assumed to contain cM. This is important for conventional VCF
data, where POS normally denotes physical base-pair position.
When a valid FORMAT column is present, columns between ALT and FORMAT
may be the VCF metadata fields QUAL, FILTER, INFO, and
the SelectionTools-specific CM field. These metadata fields must be
unique. Every column after FORMAT is a sample column, irrespective of its
name. This allows sample identifiers such as POS, INFO,
CM, or FORMAT without ambiguity.
A FORMAT column is recognized only when every marker-specific definition contains exactly one non-empty GT field. Empty FORMAT fields and duplicated GT entries are invalid. The FORMAT definition may vary among markers. The GT subfield is extracted from every sample cell. Other subfields, including DP, AD, GQ, and phase-set fields, are ignored and are not retained. A non-missing sample field that is too short to contain the GT subfield defined by FORMAT, or that contains an empty GT subfield, is rejected as malformed rather than being converted to missing. If FORMAT is absent, columns after the first five fixed fields are direct GT sample columns, except that an alphanumeric column named CM whose complete contents are valid finite nonnegative numeric values is interpreted as the genetic CM field. This rule is the convention used to distinguish an optional CM field from a sample that itself is named CM.
Genetic positions are stored as R double-precision numeric values without intentional rounding. Factor CM or POS columns are converted through their printed character values before numeric conversion, so factor level numbers are never used as map positions. Positions must be finite and nonnegative.
REF must contain exactly one non-missing allele label. It may not be
"." and may not contain a comma. ALT may be ".", denoting no
defined alternative allele, or a comma-separated list. Every ALT entry must
be non-empty, may not be ".", must differ from REF, and must be unique
within the marker. Thus definitions such as REF=A with ALT=A or ALT=C,C are
rejected.
Multiallelic markers are accepted. VCF allele index 0 denotes REF, index 1 the first ALT allele, index 2 the second ALT allele, and so on. Every observed non-missing allele index must be present in the corresponding REF/ALT definition. At most 253 ALT alleles may be defined for a marker. Together with REF this corresponds to at most 254 biological allele states in a SelectionTools marker data set; missing data use a separate internal state.
Only diploid genotypes are supported. Both "/" and "|" are
accepted. The supplied allele order is retained in allele1 and
allele2, but the separator itself is not stored. Retaining allele order
is not a statistical phasing procedure and must not be interpreted as phase
inference. Phase-set boundaries and other phase metadata are not represented
in STvcf.
Accepted complete-missing representations include R NA, an empty
field, ".", "./.", ".|.", "-1",
"-1/-1", "-1|-1", and the character value "NA".
The -1 spelling is accepted as a SelectionTools backward-compatibility
extension; standard VCF output continues to use "." for missing
alleles. Partially missing diploid genotypes such as 0/., ./1,
0|., .|1, 0/-1, and -1/1 are normalized to
complete missingness. The non-missing member of a partial pair must
nevertheless be a syntactically valid nonnegative integer allele index within
the REF/ALT definition.
Malformed fields such as an empty allele component or a nonnumeric
non-missing component are rejected rather than being hidden by the
partial-missing rule.
The converter writes allele1 and allele2 as raw matrices.
Valid non-missing VCF allele indices are 0 through 253, and raw value 255
denotes missing data. Integer allele matrices are also accepted by the STvcf
loader and reverse converter for manually constructed STvcf objects, but the
253-ALT contract still applies.
The returned list has class "STvcf" and components CHROM,
POS, ID, REF, ALT, individual,
allele1, allele2, and CHROM.levels. Marker and
individual order are retained. Marker ordering is changed only when the STvcf
object is installed in the internal SelectionTools linkage-map structure.
The conversion is not a byte-for-byte VCF round trip. QUAL, FILTER, INFO, FORMAT subfields other than GT, phase-set information, physical POS when CM is supplied, and the distinction between slash and vertical-bar separators are not stored.
Value
A list of class "STvcf" containing marker metadata, individual identifiers, two compact allele matrices, and the chromosome-label mapping. Positions are stored as double-precision values.
Construct a path in the SelectionTools data directory
Description
Construct a path in the SelectionTools data directory.
Usage
st.dd(filename)
Arguments
filename |
File name. |
Value
A character scalar containing the constructed path in the SelectionTools data directory.
st.def.hblocks()
Description
Building of haplo blocks either based on number of markers, based on a given genetic distance in cM, or for dissecting an introgression library in disjunct segments.
Usage
st.def.hblocks(hap = 5, hap.unit = 1, ld.threshold = 0.8,
ld.criterion = "none", tolerance = 0, hap.symbol = "b",
out.filename = "blocks", auxfiles = TRUE, data.set = "default")
Arguments
hap |
Defines either a number of markers (if hap.unit=1) or a genetic distance in cM (if hap.unit=2) that should be used for building haplo blocks |
hap.unit |
Set to 1 for building haplo blocks based on number of markers, set to 2 for building haplo blocks based on markers lying in a specified stretch of the chromosome (in cM). Set to 0 to build disjunct chromosome segments in an introgression population |
ld.threshold |
Function argument. |
ld.criterion |
Function argument. |
tolerance |
Function argument. |
hap.symbol |
Function argument. |
out.filename |
Character string retained for compatibility and used as part of the name of an internal file in the R session temporary directory. The temporary file is removed before the function returns. |
auxfiles |
Logical. If |
data.set |
character. Name of the file that is used for building haplo blocks |
Value
Invisibly returns a data.frame with columns Chrom, Pos, Name, Class, and Markers describing the defined haplotype blocks when auxfiles = TRUE and the operation succeeds; otherwise NULL.
st.genetic.distances()
Description
The function st.genetic.distances calculates genetic distances for each possible pair of individuals.
Usage
st.genetic.distances(measure = "mrd", format = "l",
filename = NULL, auxfiles = FALSE,
data.set = "default")
Arguments
measure |
character. Defines the measure to be used for calculation. Incorporated measures are modified Rogers distance ("mrd"), Rogers distance ("rd") and Euclidean distance ("euc"). |
format |
character. Specifies the format of the output file. Set to "l" for long format or to "m" for matrix format. |
filename |
Character string giving the base name of the genetic-distance file. Required when |
auxfiles |
Logical. If |
data.set |
character. Name of the data set that is used by the function. |
Details
The Euclidean distance is calculated as:
D_E=\sqrt{\sum_{i=1}^m\sum_{j=1}^{n_j} (p_{ij}-q_{ij})^2}
where p_{ij} and q_{ij} are allele frequencies of the jth
allele at the ith locus in the two individuals under consideration, n_i
is the number of alleles at the ith locus, and m refers to the number of loci.
The D_E ranges from zero to 2m, the limits being assumed when the
two individuals have identical allele frequencies or are fixed for different alleles.
The D_E is appropriate if allelic informative marker data are available
and the relationships between individuals are investigated in combination
with multivariate methods that require dissimilarities possessing the Euclidean property (Reif 2005).
The Rogers distance (Rogers 1972) is a modification of D_E and is calculated as follows:
D_R=\frac{1}{m}\sum_{i=1}^m \sqrt{\frac{1}{2}\sum_{j=1}^{n_j} (p_{ij}-q_{ij})^2}
The D_R is the average D_E across all loci standardized with the factor 1/2
to restrict the values to the interval [0,1]. It is one only if two individuals
are fixed for different alleles, but if one or both individuals are not fixed
and they have no alleles in common, D_R is not equal to one.
Melchinger et al. (1991) derived theoretical results that D_R estimates
between two homozygous inbreds are linearly related to the coancestry coefficient.
Consequently, D_R is suitable for studying the relationship between
the genetic dissimilarity of inbreds based on allelic informative marker
data and the coefficient of coancestry.
The modified Rogers distance (Wright 1978) is calculated with the formula:
D_W=\frac{1}{\sqrt{2m}}\sqrt{\sum_{i=1}^m\sum_{j=1}^{n_j} (p_{ij}-q_{ij})^2}
As an Euclidean distance with the values in [0,1] it can be used for the
same applications as recommended for D_E. Like D_R, D_W
is not equal to one in the case of multiple alleles, even if the two individuals
have no allele in common.
D_W is especially suitable in studies based on allelic informative
marker data for examining the prediction of heterosis with genetic dissimilarities
or the establishment of heterotic groups. Furthermore, D_W can be used
for the same applications as suggested for D_E (e.g., multivariate
methods), owing to its Euclidean property.
The Rogers distance and the modified Rogers distance are
standardized measures in the interval [0,1]. The standardization is
achieved by the divisions by m and by 2 in the respective equations.
This requests codominant markers! These measures can't be used with
dominant markers. However the Euclidean distance is not standardized and
can also be used with dominant marker data.
For homozygous genotypes Rogers distance equals the Nei-Li Distance
(Nei and Li 1979), which is the same as 1 minus the Dice coefficient
(Dice 1945), further the modified Rogers distance is the square root of the Rogers distance.
Value
Invisibly returns either an object of class "dist" when format = "m", representing pairwise genetic distances, or a long-format data.frame with columns OTU1, OTU2, and Measure when format = "l".
References
Dice, L.R. 1945. Measures of the amount of ecologic association between species. Ecology 26:197–302.
Nei, M. and W.H. Li. 1979. Mathematical Models for studying genetic variation in terms of restriction endonucleases. Proc. Natl. Acad. Sci. USA. 76:5268–5371.
Reif, J., Melchinger, A. and Frisch, M. 2005. Genetical and mathematical properties of similarity and dissimilarity coefficients applied in plant breeding and seed bank management. Crop Sci.45:1-7.
Rogers, J.S. 1972. Measures of genetic similarity and genetic distance. VII. Univ. Tex. Publ. 7213:145–153.
Wright, S. 1978. Evolution and the Genetics of Populations. Volume 4: Variability Within and Among Natural Populations. University of Chicago Press, Chicago, Illinois.
Calculate genetic distances from SelectionTools marker data
Description
Calculate genetic distances from SelectionTools marker data.
Usage
st.genetic.distances.02(measure = "mrd", format = "l",
filename = NULL, auxfiles = FALSE,
data.set = "default")
Arguments
measure |
Distance or similarity measure. |
format |
Input or output format. |
filename |
Character string giving the base name of the genetic-distance file. Required when |
auxfiles |
Logical. If |
data.set |
Name of the SelectionTools data set. |
Value
Invisibly returns either an object of class "dist" when format = "m", representing pairwise genetic distances, or a long-format data.frame with columns OTU1, OTU2, and Measure when format = "l".
Calculate genetic distances from SelectionTools marker data
Description
Calculate genetic distances from SelectionTools marker data.
Usage
st.genetic.distances.fct(measure = "mrd", format = "l", split = 1,
filename = NULL, auxfiles = FALSE,
data.set = "default")
Arguments
measure |
Distance or similarity measure. |
format |
Input or output format. |
split |
Split or partition control. |
filename |
Character string giving the base name of the genetic-distance file. Required when |
auxfiles |
Logical. If |
data.set |
Name of the SelectionTools data set. |
Value
Invisibly returns either an object of class "dist" when format = "m", representing pairwise genetic distances, or a long-format data.frame with columns OTU1, OTU2, and Measure when format = "l".
Return the current SelectionTools information level
Description
Return the current SelectionTools information level.
Usage
st.get.info.level()
Value
An integer scalar giving the current SelectionTools information level.
Return the linkage map for a SelectionTools data set
Description
Return the linkage map for a SelectionTools data set.
Usage
st.get.map(filename = "map", data.set = "default")
Arguments
filename |
Character string retained for compatibility and used as part of the name of the internal session-temporary transfer file. No output file is retained. |
data.set |
Name of the SelectionTools data set. |
Value
Invisibly returns a data.frame with columns Chrom, Pos, Name, and Class describing the map of data.set; returns NULL if no map can be returned.
Return the configured number of computational threads
Description
Return the configured number of computational threads.
Usage
st.get.num.threads()
Value
An integer scalar giving the configured number of computational threads.
Return Simulation-Population Data in SelectionTools Format
Description
Imports a diploid simulation population into a SelectionTools marker data set.
Usage
st.get.simpop(pop.name, data.set ="default")
Arguments
pop.name |
Character string naming the simulation population to import. For this transfer the population name must contain 2 to 255 ASCII letters or digits. |
data.set |
Character string naming the SelectionTools marker data set to create or replace. |
Details
The simulation population is converted to the two-allele marker matrix representation used by SelectionTools. The simulation genome must be diploid. Marker identities and loci are obtained from the active simulation map. Marker names transferred from that map must contain only ASCII letters and digits; invalid names are rejected rather than modified. Simulation missing alleles are converted to the SelectionTools internal missing representation, and marker statistics are calculated for the imported matrix. Non-missing simulation allele codes must lie between 1 and 254 to be transferred into a SelectionTools marker data set; wider simulation allele codes are rejected rather than narrowed or wrapped.
The transfer creates new marker-data individual identifiers rather than recovering names that may have existed before the population entered simulation. For a population named P1, the generated identifiers are P1I1, P1I2, and so on. The separator I is alphanumeric so the resulting names obey the marker-data/STvcf convention that individual identifiers contain only letters and digits. Population names containing dots, punctuation, whitespace, or other non-alphanumeric characters are rejected by this transfer rather than being silently sanitized.
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of copying a simulation population into the named SelectionTools marker data set.
Return simulation-population data in SelectionTools format
Description
Imports a simulation population as marker data and uses an evaluated simulation effect as the phenotype vector of the SelectionTools data set.
Usage
st.get.simpop.perfdata(pop.name, data.set, effect)
Arguments
pop.name |
Character string naming the simulation population. |
data.set |
Character string naming the target SelectionTools marker data set. |
effect |
Effect name or effect specification passed to population evaluation. |
Details
The function first imports the simulation population into data.set. It then ensures that the simulation population is genotyped, evaluates the requested effect, retrieves the resulting population genetic values, and writes those values into the phenotype vector of the imported SelectionTools marker data set.
The operation therefore combines a simulation-to-marker-data transfer with generation of performance data from the simulation effect model. The diploidy and naming requirements of the simulation-population importer apply. In particular, the population name used by this transfer must contain only ASCII letters and digits, because st.get.simpop() constructs alphanumeric marker-data individual identifiers from that name.
Value
An invisible integer status value returned by gs.set.performance.data() after simulation-population performance data have been transferred to the SelectionTools data set.
Return Simulation-Population Marker Data without a Linkage Map
Description
Imports diploid simulation-population marker genotypes into a SelectionTools marker data set using the alternative simulation transfer path.
Usage
st.get.simpop2(pop.name, data.set ="default")
Arguments
pop.name |
Character string naming the simulation population. For this transfer the population name must contain 2 to 255 ASCII letters or digits. |
data.set |
Character string naming the SelectionTools marker data set. |
Details
The transfer is restricted to diploid simulation data. Generated individual
identifiers use the same convention as st.get.simpop(): a population
P1 produces P1I1, P1I2, and so on. Population names and marker names
transferred from the simulation data must contain only ASCII letters and
digits. Invalid names are rejected rather than modified.
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of copying simulation-population marker data without a linkage map into the named SelectionTools data set.
Construct a path in the SelectionTools input directory
Description
Construct a path in the SelectionTools input directory.
Usage
st.id(filename)
Arguments
filename |
File name. |
Value
A character scalar containing the constructed path in the SelectionTools input directory.
Construct a path in the SelectionTools input directory
Description
Construct a path in the SelectionTools input directory.
Usage
st.indir(filename)
Arguments
filename |
File name. |
Value
A character scalar containing the constructed path in the SelectionTools input directory.
Print or handle a SelectionTools information message
Description
Print or handle a SelectionTools information message.
Usage
st.info(lev, msg)
Arguments
lev |
Information or verbosity level. |
msg |
Message text. |
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of issuing the requested SelectionTools information message.
Load Performance Data from an R Data Set
Description
Matches performance records supplied in an R object to the individuals of an existing SelectionTools marker data set without intermediate text files.
Usage
st.load.performance.data(performance.data, data.set="default")
Arguments
performance.data |
A data frame or matrix containing individual identifiers and numeric performance values. |
data.set |
Character string naming the SelectionTools marker data set. |
Details
Marker data must already be present in data.set. Three input-column
conventions are accepted. Columns named individual and
performance are used when both are present; otherwise columns
ii and yy are used when both are present; otherwise an object
with exactly two columns is interpreted as identifier followed by performance.
Performance individual IDs must contain only ASCII letters and digits and fewer than 256 bytes. The same name convention is used for STvcf individual identifiers. Invalid IDs are rejected rather than being silently ignored as unmatched records.
Records are matched by individual identifier. Performance records whose valid IDs do not occur in the current marker data are ignored. Marker-data individuals with no matched performance record are removed. If an identifier occurs more than once in the performance input, the last matched value is used, following the existing performance-input convention. At least one record must match.
Performance values must be finite numeric values after conversion. The resulting marker data set retains the current marker names and linkage map, subsets both allele arrays to individuals with performance, recalculates marker statistics for the retained individuals, and stores the matched performance vector. The operation is transactional.
Value
NULL, returned invisibly. The function is called for its side effect of loading performance data into the named SelectionTools data set.
Load Compact STvcf Marker and Linkage-Map Data
Description
Loads a compact STvcf object directly into a SelectionTools marker data set and installs its linkage map without intermediate marker or map text files.
Usage
st.load.vcf.data(STvcf, data.set="default")
Arguments
STvcf |
A list in STvcf format. |
data.set |
Character string naming the SelectionTools marker data set to be created or replaced. |
Details
STvcf is a compact exchange representation for diploid marker data with VCF-style allele indexing and one genetic position per marker. Loading converts it into the existing SelectionTools marker and linkage-map data structures; no parallel internal analysis representation is introduced.
The list must contain exactly the following nine components, each exactly once. Extra, missing, duplicated, unnamed, or unknown list components are rejected by the C loader:
CHROMInteger running chromosome codes, consecutive from 1.
POSDouble-precision genetic positions in cM. Values must be finite and nonnegative.
IDUnique marker identifiers containing only ASCII letters and digits and fewer than 256 bytes.
REFOne reference-allele label per marker. REF may not be missing, empty,
".", or comma-separated.ALT"."or a comma-separated alternative-allele list. ALT entries must be non-empty, unique, and different from REF.individualUnique individual identifiers containing only ASCII letters and digits and fewer than 256 bytes.
allele1Marker-by-individual raw or integer matrix containing the first VCF allele index.
allele2Marker-by-individual raw or integer matrix containing the second VCF allele index.
CHROM.levelsCharacter labels associated with the running chromosome codes.
Name validation is repeated in C. A manually assembled STvcf therefore cannot bypass the rule that marker and individual names contain only letters and digits. Dots, underscores, hyphens, whitespace, punctuation, non-ASCII characters, empty names, duplicate names, and names of 256 bytes or more are rejected.
STvcf allele index 0 denotes REF and indices 1, 2, ... denote ALT alleles. On
loading these become SelectionTools biological allele codes 1, 2, 3, ... .
At most 253 ALT alleles may be defined at one marker, so the largest resulting
SelectionTools biological allele code is 254. The two allele matrices must use
the same storage type. Raw STvcf matrices use 255 for missing; integer matrices
use NA_integer_ for missing data. Missingness is converted to the
SelectionTools internal missing representation.
REF and ALT definitions are validated independently of the R converter. A REF value repeated in ALT, a duplicated ALT value, an empty ALT token, or a dot embedded inside an ALT list is rejected. Thus malformed manually created STvcf objects cannot create distinct internal integer alleles for identical textual allele labels.
The SelectionTools marker representation allows at most 254 biological allele codes per marker, plus the internal missing state. Multiallelic markers are valid input as long as they fit this representation. The ordinary statistical genomic selection functions that construct one design-matrix column per marker are biallelic; multiallelic data must be restricted to an appropriate biallelic marker set before those functions are called. Monomorphic markers are also unsuitable for that ordinary biallelic design-matrix construction.
Only diploid data are supported. If either stored allele is missing, the complete internal diploid genotype is set to missing. The order of two non-missing alleles is retained, and those two arrays become the two homologues when marker data are transferred to simulation. This does not imply statistical phase inference.
All STvcf markers require map positions. Markers are loaded in chromosome and
cM order. Exact position ties retain STvcf row order and are separated by
small deterministic internal offsets because the existing linkage-map and
simulation structures require distinct ordered positions. The STvcf object
itself is not changed; st.get.map() reports the internally used map
positions.
Normal SelectionTools marker statistics are constructed after conversion. They are calculated from exactly the markers present in the STvcf object. Compact raw or integer R storage is expanded into the existing C genotype and marker-statistics structures during loading.
The operation replaces data.set transactionally. The replacement is
committed only after names, alleles, genotype matrices, linkage-map data, and
marker statistics have been validated and constructed successfully. On
failure, the previous contents of the named marker data set remain intact.
Performance data are not installed by this function.
Value
NULL, returned invisibly. The function is called for its side effect of loading the supplied STvcf marker and map information into data.set.
Identify markers by effect and replace selected marker alleles in marker data
Description
Identify markers by effect and replace selected marker alleles in marker data.
Usage
st.mark.alleles(markers = NULL, data.set="default", effect.set=NULL,
target.set, hap.lst = NULL, alpha=0.05,
p.adjust.method="none", replace = "3")
Arguments
markers |
Function argument used by this operation. |
data.set |
Name of the SelectionTools data set. |
effect.set |
Data set supplying marker effects. |
target.set |
Target data set. |
hap.lst |
Function argument used by this operation. |
alpha |
Method-specific tuning or significance parameter. |
p.adjust.method |
Function argument used by this operation. |
replace |
Logical or coding value controlling replacement. |
Value
On successful completion, invisibly returns the character value used to restore the previous SelectionTools input directory. Validation failures return NULL invisibly. The main effect is the requested allele recoding and marker-data update.
st.marker.data.statistics()
Description
Recomputes and optionally returns summary statistics for the current marker matrix.
Usage
st.marker.data.statistics(data.set ="default", filename ="marker.stats",
ind = TRUE, mar = TRUE, gen = TRUE, auxfiles = TRUE)
Arguments
data.set |
Character string naming the marker data set. |
filename |
Character string retained for compatibility and used as part of the names of internal statistics files in the R session temporary directory. These temporary files are removed before the function returns. |
ind |
Logical. Request the individual missingness table. |
mar |
Logical. Request the marker statistics table. |
gen |
Logical. Request the genotype matrix table. |
auxfiles |
Logical. If |
Details
Statistics are calculated directly from the two stored allele matrices. For each marker the function determines the distinct biological allele codes, allele counts, the proportion of missing allele copies, and expected heterozygosity calculated as one minus the sum of squared non-missing allele frequencies. The internal missing-allele sentinel is excluded from the allele-frequency denominator.
For each individual, the proportion of markers with a missing first allele is reported as the individual missingness statistic. These statistics are also the values used by marker-data restriction functions.
The marker summary distinguishes stored missingness from biological allele codes. Missing data are not counted as a biological allele for restrictions such as NoAll.MAX. Biological allele codes are in the range 1 through 254.
When requested, three files are produced: an individual summary, a marker summary, and a genotype matrix. The R wrapper reads these files back and removes the temporary files.
Value
Invisibly returns a named list containing the requested marker-data summaries. Depending on ind, mar, and gen, components are individual.list, marker.list, and/or genotypes, each a data.frame; NULL is returned if no result can be produced.
Convert a SelectionTools Marker Data Set to STvcf
Description
Converts marker genotypes and the linkage map of an existing SelectionTools marker data set to the compact diploid STvcf representation.
Usage
st.markerdata.to.STvcf(data.set="default")
Arguments
data.set |
Character string naming the SelectionTools marker data set to be converted. |
Details
This function supplies the marker-data to STvcf step of the reverse data
chain. With the existing st.get.simpop() function, a simulation
population can first be returned to the ordinary SelectionTools marker-data
structure and can then be converted to STvcf without introducing a separate
simulation-to-STvcf bridge.
The marker data set must contain genotype data and a usable linkage map. Every marker-matrix column must have exactly one active map location. Map-only locations are ignored. Missing or duplicate marker locations, inconsistent marker and map names, invalid chromosome numbers, non-finite positions, and negative positions cause conversion to fail.
The current marker order is retained. Map positions are returned as the
current double-precision cM values without intentional rounding. Current
numeric chromosome numbers are converted to consecutive STvcf running codes
1, 2, ..., in order of first appearance, while CHROM.levels records
the corresponding numeric chromosome labels as character strings.
Current marker and individual names become STvcf identifiers. They must be unique, contain fewer than 256 bytes, and contain only ASCII letters and digits. Dots and all other punctuation are rejected. Names are never silently modified during this conversion.
When the marker data came from a simulation population through
st.get.simpop(), that existing transfer deliberately creates new
individual identifiers. In the revised convention a population named P1
produces names P1I1, P1I2, and so on. The population name used by
st.get.simpop() must itself contain only letters and digits so that
the generated individual identifiers satisfy the marker-data/STvcf name
rules. Original individual names that existed before simulation are not
reconstructed; this name change is an intended feature of the current
simulation-to-marker-data transfer.
The SelectionTools marker structure stores integer allele states but not the
original textual VCF REF/ALT labels. The converter therefore normalizes the
observed non-missing allele states separately for each marker. States are
ordered by their current integer values. The first state becomes
REF="1"; further states become ALT="2,3,..."; genotypes are
converted to VCF indices 0, 1, 2, ... in that order.
For marker data originally loaded from STvcf, internal alleles are already 1 for REF, 2 for ALT1, 3 for ALT2, and so on. Transfer to and from simulation preserves those integer states. For marker data created by other routes, distinct genotype states are retained but historical integer labels are normalized.
An all-missing marker is retained with REF="1", ALT=".", and
missing genotypes. A marker with one observed non-missing state is likewise
written with ALT="." and uses VCF allele index 0. If either homologue
of a marker-data genotype is missing, the exported STvcf genotype is
completely missing, consistent with the STvcf partial-missing convention.
The two current marker-data allele arrays are retained in order as
allele1 and allele2. For data obtained through
st.get.simpop(), they correspond to the two simulation homologues. No
new phasing calculation is performed.
The current SelectionTools marker representation permits at most 254 biological allele states per marker. After normalization their VCF indices are therefore at most 253, so this converter writes raw STvcf allele matrices with raw value 255 for missing. Marker genotypes that violate the SelectionTools allele-range or missingness invariants are rejected.
The linkage-map class is not part of STvcf. Exact map-position ties may have already been dispersed by the SelectionTools linkage-map or simulation machinery; this function returns the current stored positions and cannot reconstruct an earlier duplicate position that is no longer represented.
Only marker genotype and genetic-map information are returned. Performance data, estimated effects, genomic relationship matrices, selection values, simulation GValue/PValue arrays, and other derived objects are outside STvcf.
Value
A list of class "STvcf" containing the current marker data, individual identifiers, marker metadata and chromosome-label mapping. If conversion fails, NULL is returned.
Fit the package multi-kernel mixed-model routine
Description
Fit the package multi-kernel mixed-model routine.
Usage
st.mxd(y, X, Z = NULL, V = NULL, estimates = FALSE, maxcyc = 100,
precision = 1e-4)
Arguments
y |
Response vector. |
X |
Fixed-effect design matrix. |
Z |
Marker or design matrix. |
V |
Function argument used by this operation. |
estimates |
Function argument used by this operation. |
maxcyc |
Maximum number of fitting cycles. |
precision |
Convergence tolerance. |
Value
A named list on successful fitting, with components:
sigma |
Numeric vector of variance-component estimates. |
b |
Fixed-effect estimates when |
u |
Random-effect estimates for |
Returns NULL when input validation or the compiled fit reports failure.
Construct a path in the SelectionTools output directory
Description
Construct a path in the SelectionTools output directory.
Usage
st.od(filename)
Arguments
filename |
File name. |
Value
A character scalar containing the constructed path in the SelectionTools output directory.
Construct a path in the SelectionTools output directory
Description
Construct a path in the SelectionTools output directory.
Usage
st.outdir(filename)
Arguments
filename |
File name. |
Value
A character scalar containing the constructed path in the SelectionTools output directory.
Plot correlation information from SelectionTools results
Description
Plot correlation information from SelectionTools results.
Usage
st.plot.corr(x, y, title = "", pch = 16, cex = 1, xlab = "x", ylab = "y")
Arguments
x |
Input object. |
y |
Response vector. |
title |
Function argument used by this operation. |
pch |
Function argument used by this operation. |
cex |
Function argument used by this operation. |
xlab |
Function argument used by this operation. |
ylab |
Function argument used by this operation. |
Value
No useful return value is intended. The function is called for its side effect of plotting correlation information.
Plot correlation information from SelectionTools results
Description
Plot correlation information from SelectionTools results.
Usage
st.plot.corr.l(x, y, title = "", pch = 16, cex = 1, xlab = "x", ylab = "y",
ylim, xlim)
Arguments
x |
Input object. |
y |
Response vector. |
title |
Function argument used by this operation. |
pch |
Function argument used by this operation. |
cex |
Function argument used by this operation. |
xlab |
Function argument used by this operation. |
ylab |
Function argument used by this operation. |
ylim |
Function argument used by this operation. |
xlim |
Function argument used by this operation. |
Value
No useful return value is intended. The function is called for its side effect of plotting correlation information.
st.plot.gene.diversity()
Description
The function st.plot.gene.diversity plots gene diversity along the chromosomes and returns the marker-level values used for the plot.
Usage
st.plot.gene.diversity(data.set ="default", plt.pdf = FALSE,
plt.fname = NULL, plt.width = 10, plt.height = 6,
plt.ptsize = 14)
Arguments
data.set |
character. Name of the data set that is used by the function. |
plt.pdf |
logical. If |
plt.fname |
Character string giving the base name of the PDF file. Required when |
plt.width |
Width of the generated PDF file. |
plt.height |
Height of the generated PDF file. |
plt.ptsize |
Point size of the generated PDF file. |
Details
The gene diversity (expected heterozygosity) at a marker is calculated as
1-\sum_i p_i^2,
where p_i is the frequency of allele i at that marker. The values
used for plotting are obtained from st.marker.data.statistics(). The
average gene diversity for each chromosome is printed at the top of the plot.
Value
Invisibly returns a data.frame with columns chrom, pos, name, cpos, and gene.div, containing the marker-level positions and gene-diversity values used for plotting.
st.plot.ggt()
Description
The function st.plot.ggt plots graphical genotypes for all or a subset of individuals.
Usage
st.plot.ggt(data.set = "default", ifilename = "", i.list = "", f.ind = 0,
l.ind = 0, f.lr = 2, f.tb = 2, d.h = 0.001, d.v = 0.001, p.t = FALSE,
p.s = FALSE, d.t = 50, d.map = 100, c.nme = "", i.nme = "",
cex.chrom = 1, cex.ind = 1,
color = c ("yellow","blue","red","green", "wheat",
"skyblue","tomato","palegreen",
"yellow4","darkblue",
"darkblue","darkgreen" ),
z.min = 0, z.max = 12, plt.pdf = FALSE, plt.png = FALSE,
plt.fname = NULL, plt.width = 10, plt.height = 10,
plt.ptsize = 10)
Arguments
data.set |
character. Uses the data set that was created by st.read.marker.data. If no data set was specified, the "default" data set is used. |
ifilename |
Name of a file containing a list of individuals of which the graphical genotypes should be plotted. The file is by default read from the input directory. |
i.list |
Function argument. |
f.ind |
Position of first individual used for plotting. |
l.ind |
Position of last individual used for plotting. |
f.lr |
Width between color bars and labels (individual names). |
f.tb |
Height between color bars and top and bottom labels (chromosome names and marker names). |
d.h |
Horizontal distance between chromosomes. |
d.v |
Vertical distance between individuals. |
p.t |
logical. If TRUE marker tickmarks are plotted. |
p.s |
logical. If TRUE, a cM scale is plotted. |
d.t |
Length of marker tickmarks. |
d.map |
Distance between marker labels. |
c.nme |
Character string specifying the names of chromosomes. |
i.nme |
Character string specifyin the names of individuals. |
cex.chrom |
Character expansion factor for chromosome names. |
cex.ind |
Character expansion factor for individual names. |
color |
Character vector specifying the colors used to represent genotype or allele codes. |
z.min |
Minimum genotype or allele code represented by the color scale. |
z.max |
Maximum genotype or allele code represented by the color scale. |
plt.pdf |
logical. If TRUE a pdf file is generated. |
plt.png |
logical. If TRUE, a PNG file is generated. |
plt.fname |
Character string giving the base name of the output file. Required when |
plt.width |
Width of generated pdf file. |
plt.height |
Height of generated pdf file. |
plt.ptsize |
Point size of generated pdf file. |
Details
The color coding for SNP data is by default A=yellow, C=blue, G=red, T=green. For
haplo blocks the coding is dependent on the frequency of a particular haplo block.
The most frequent haplo block is plotted in yellow, the second frequent in blue
and so on.
Using the function st.plot.ggt after the function st.recode.hbc gives a good graphical
overview, which haplo blocks are not inherited from the reference genotype (if reference
individual = 1, the reference parent genome is plotted in yellow. The haplo
blocks inherited from the donor parent are plotted in blue.
Using the function st.plot.ggt after the function st.recode.hil gives a good graphical
overview, which haplo blocks are inherited from which parent. For plotting, the individuals are sorted
by their haplotype frequency. Thus, the parents are usually plotted on top in yellow and blue (default values).
The haplo blocks of the individuals are also plotted in yellow or blue, if they are completely identical
to a parent's haplo block. If they are plotted in a different color, a cross-over did occur and thus,
this haplo block is not identical one of the parent's haplo blocks.
If an external list of individuals is used, the ordering of individuals
within that list is used for plotting. When plotting on the current
graphics device, the function stops if that device is already in
split-screen mode, so that an existing user-defined split-screen layout
is not changed.
Value
On a successful run, an invisible implementation-level list returned by the final timing/information helper; the meaningful result is the side effect of producing the graphical genotype plot. Validation failures may return NULL.
Create a graphical genotype plot from SelectionTools source data
Description
Create a graphical genotype plot from SelectionTools source data.
Usage
st.plot.ggt.src(map, pop, f.lr, f.tb, d.h, d.v, p.t, p.s, d.t, d.map, c.nme,
i.nme, color, z.min, z.max, cex.chrom, cex.ind, plt.png,
plt.pdf, plt.name, plt.width, plt.height, plt.ptsize)
Arguments
map |
Linkage map or map object. |
pop |
Population name or population specification. |
f.lr |
Function argument used by this operation. |
f.tb |
Function argument used by this operation. |
d.h |
Function argument used by this operation. |
d.v |
Function argument used by this operation. |
p.t |
Function argument used by this operation. |
p.s |
Function argument used by this operation. |
d.t |
Function argument used by this operation. |
d.map |
Map-related parameter. |
c.nme |
Function argument used by this operation. |
i.nme |
Function argument used by this operation. |
color |
Function argument used by this operation. |
z.min |
Function argument used by this operation. |
z.max |
Function argument used by this operation. |
cex.chrom |
Chromosome-related parameter. |
cex.ind |
Function argument used by this operation. |
plt.png |
Function argument used by this operation. |
plt.pdf |
Function argument used by this operation. |
plt.name |
Function argument used by this operation. |
plt.width |
Function argument used by this operation. |
plt.height |
Function argument used by this operation. |
plt.ptsize |
Function argument used by this operation. |
Value
No useful return value is intended. The function is called for its side effect of producing a graphical genotype plot and, when requested, closing the graphics device.
st.read.map()
Description
Reads a linkage map for an existing marker data set, matches loci by marker name, and restricts the marker matrix to markers represented in both sources.
Usage
st.read.map(filename, skip = 0, format = "cpms", m.stretch = 1,
auxfiles = TRUE, data.set = "default")
Arguments
filename |
Character string naming the map input file. Absolute paths are used directly; relative names are resolved using |
skip |
Non-negative integer giving the number of initial input lines to ignore. |
format |
Character string specifying |
m.stretch |
Finite numeric multiplier applied to every map position after reading. Use 1 when positions already use the desired map unit. |
auxfiles |
Logical. If |
data.set |
Character string naming the marker data set to which the map is attached. |
Details
A marker data set must already be loaded. The supported map layouts are "cpms", containing chromosome, position, marker name, and class, and "mcp", containing marker name, chromosome, and position. In "mcp" input the class is assigned internally. The first skip lines can be ignored before map parsing. Map positions are multiplied by m.stretch.
Marker matching is by marker name. Loci present in the map but absent from the marker matrix are silently ignored. Marker rows present in the marker matrix but absent from the map are silently removed. This behavior is intentional: after a successful read the marker data set represents the intersection of marker matrix and linkage map. The remaining genotype rows are reordered into linkage-map order and marker statistics are recalculated. Phenotypic data associated with individuals are preserved.
Duplicate marker names in either the linkage map or marker matrix are rejected because they make the mapping ambiguous. Equal positions on the same chromosome are separated deterministically while preserving their relative order and without moving a tied locus beyond the next genuinely distinct map position.
The operation is transactional. Parsing, matching, construction of the restricted genotype matrix, linkage-map structures, and marker statistics are completed before the existing marker data are replaced. On failure, the previous marker matrix, map, phenotype vector, and derived state remain available. On successful replacement, model- and marker-order-dependent derived structures belonging to the old marker set are invalidated.
Value
Invisibly returns a data.frame with columns Chrom, Pos, Name, and Class when auxfiles = TRUE and map import succeeds; otherwise NULL. The main effect is updating the map associated with data.set.
st.read.marker.data()
Description
Reads marker data into a named SelectionTools marker data set and derives the marker statistics used by subsequent analyses.
Usage
st.read.marker.data(filename, format, auxfiles = FALSE, data.set = "default")
Arguments
filename |
Character string naming the marker input file. Absolute paths are used directly; relative names are resolved using |
format |
Character string selecting the input layout: |
auxfiles |
Logical. If |
data.set |
Character string naming the SelectionTools marker data set that is replaced after a successful read. |
Details
Four input layouts are supported. Format "l" is a list or long format in which a record identifies an individual, a marker, and one allele. A homozygous genotype can therefore be represented by one allele record and a heterozygous genotype by two records for the same individual and marker. In this format a multi-digit token such as 12 is a single allele number 12.
Format "m" is a marker-major matrix. Individual names are read from the header and each subsequent row starts with a marker name followed by one genotype for each individual. Format "t" is the transposed, individual-major matrix and uses semicolon-separated fields. Format "n" is the NTSys/Plabsim allele-incidence representation in which rows have names of the form marker.allele followed by zero-one incidence values.
In matrix genotype fields, a single allele is interpreted as homozygous. A slash separates explicitly coded alleles, such as 12/14. Compact two-character genotypes are interpreted as two one-character alleles, so 12 denotes genotype 1/2. Compact nucleotide genotypes are also recognized; the internal nucleotide codes are A=1, C=2, G=3, T=4, and D=5. Numeric biological allele codes must be integers from 1 through 254; they do not need to be consecutive within a marker. The tokens -1, --, -, ., and NA denote missing data. The -1 spelling is retained for compatibility with historical SelectionTools marker files. If either allele of a diploid genotype is missing, the complete genotype is stored as missing.
The read is transactional. The new marker data, names, genotype matrix, and marker statistics are built and validated before replacing an existing data set. If the file cannot be read or is structurally invalid, the previous marker data set is left unchanged. Duplicate marker or individual identities that would make a matrix ambiguous are rejected. NTSys marker blocks must be contiguous and incidence values must be zero or one.
After a successful read the marker statistics are recomputed from the imported matrix. When auxfiles=TRUE, the R wrapper reads the generated individual, marker, and genotype summaries back into R and removes the temporary output files.
Value
Invisibly returns a list when the requested auxiliary tables are produced, with components individual.list, marker.list, and genotypes, each a data.frame. Otherwise the return value is NULL; the main effect is loading marker data into data.set.
Read marker data from a data frame into SelectionTools
Description
Read marker data from a data frame into SelectionTools.
Usage
st.read.marker.data.df(markerdata, format = NULL, data.set = "default")
Arguments
markerdata |
Function argument used by this operation. |
format |
Input or output format. |
data.set |
Name of the SelectionTools data set. |
Details
Biological allele codes must be integers from 1 through 254; they need not be
consecutive within a marker. R missing allele values and the character tokens
-1, --, -, ., and NA are treated as missing.
The -1 spelling is retained for compatibility with historical
SelectionTools marker data. If
either allele of a diploid genotype is missing, the complete genotype is
stored as missing. Repeated individual/marker records are aggregated
deterministically using the same rules as list-format marker input.
Value
NULL, returned invisibly. The function is called for its side effect of loading the supplied marker-data frame into the named SelectionTools data set.
st.read.performance.data()
Description
Reads phenotypic values for a marker data set and retains only marker-data individuals for which a phenotype is available.
Usage
st.read.performance.data(filename, out.filename = "y.vector", auxfiles = TRUE,
data.set ="default")
Arguments
filename |
Character string naming the phenotype input file. Relative names are resolved using |
out.filename |
Character string retained for compatibility and used as part of the name of an internal file in the R session temporary directory. The temporary file is removed before the function returns. |
auxfiles |
Logical. If |
data.set |
Character string naming the marker data set whose individuals are matched to the phenotype file. |
Details
The input contains individual identifiers and phenotypic values. The marker matrix must already be loaded because individual names are matched against the marker-data individuals.
Phenotype records for individuals that are not present in the marker matrix are silently ignored. This is a normal situation and does not generate a warning. Conversely, individuals present in the marker matrix for which no phenotypic record is found are removed from the marker data set. Thus a successful call leaves the individual intersection of marker and phenotype data. Marker order and the linkage map are preserved, while marker statistics are recalculated for the retained individuals.
If a known individual occurs more than once in the phenotype input, the last record is used. A phenotype for a known individual must be a valid finite numeric value. Unknown individuals are ignored before their value field is interpreted. If no phenotype record matches any marker-data individual, the operation fails rather than creating an empty marker data set.
The operation is transactional. The reduced marker matrix, names, phenotype vector, map copy, and marker statistics are prepared before replacing the existing data set. A read or allocation failure therefore leaves the previous marker data set unchanged.
Value
Invisibly returns a data.frame with columns i (individual identifier) and y (performance value) when auxfiles = TRUE and the operation succeeds; otherwise NULL. The main effect is updating performance data in data.set.
st.recode.hbc()
Description
The function st.recode.hbc recodes the haplotype data for backcross populations according to a given reference individual.
Usage
st.recode.hbc(reference = 1, data.set = "default")
Arguments
reference |
Specifies which individual is treated as reference for recoding. Usually the recurrent parent is set to be the reference individual. Must be between 1 and the number of individuals in the data set. |
data.set |
Name of data set that is used for recoding. |
Details
If, at a given locus, an individual carries a haplotype that is also
present in the reference genotype, then this haplotype is recoded to 1.
If the haplotype is not present in the reference genotype, then it is
recoded to 2.
Using the function st.plot.ggt after the function st.recode.hbc gives a good graphical
overview, which haplo blocks are not inherited from the reference genotype (if reference
individual = 1, the reference parent genome is plotted in yellow. The haplo
blocks inherited from the donor parent are plotted in blue.
Value
On a successful run, an invisible implementation-level list returned by the final timing/information helper; the meaningful result is the side effect of recoding marker data according to the HBC reference convention. Validation failures may return NULL.
st.recode.hil()
Description
The function st.recode.hil recodes the haplotype data for inbred populations.
Usage
st.recode.hil(out.filename = "blocks", auxfiles = TRUE, data.set = "default")
Arguments
out.filename |
Character string retained for compatibility and used as part of the name of an internal file in the R session temporary directory. The temporary file is removed before the function returns. |
auxfiles |
Logical. If |
data.set |
Name of data set that is used for recoding. |
Details
The haplo block i of individual j at locus m is compared to the same haplo block
of the other individuals. If the haplo blocks of different individuals are identical
they get the same number. In extreme, there are as many different haplo blocks as individuals for
a given locus. If a haplotype block would result in more than 254 distinct haplotype alleles,
recoding stops with an error rather than merging or truncating haplotypes.
Using the function st.plot.ggt after the function st.recode.hil gives a good graphical
overview, which haplo blocks are inherited from which parent. For plotting, the individuals are sorted
by their haplotype frequency. Thus, the parents are usually plotted on top in yellow and blue (default values).
The haplo blocks of the individuals are also plotted in yellow or blue, if they are completely identical
to a parent's haplo block. If they are plotted in a different color, a cross-over did occur and thus,
this haplo block is not identical one of the parent's haplo blocks.
Value
Invisibly returns a data.frame with columns Block, AlleleNr, and AlleleDef describing recoded haplotype-block alleles when auxfiles = TRUE and the operation succeeds; otherwise NULL.
st.recode.ref()
Description
Compares the alleles of an individual with the alleles for a reference individual.
Usage
st.recode.ref(reference = 0, descending = FALSE, data.set = "default")
Arguments
reference |
Specifies which individual is the reference for recoding |
descending |
Check whether the alleles can originate from the reference indivdiual. Default: FALSE |
data.set |
Name of data set that is used for recoding |
Details
descending=FALSE (default):If an individuals has two alleles in
common with the reference, then it gets the genotype 1/1, for one allele
in common 1/2, and for non allele in common 2/2. The reference
individual is recoded to 1/1. Missing source genotypes remain missing.
descending=TRUE: If both alleles of an individual can originate
from the referende: 1/1, if one allele can originate from the reference:
1/2, if no allele can originate from the reference 2/2.
Value
On a successful run, an invisible implementation-level list returned by the final timing/information helper; the meaningful result is the side effect of recoding marker data according to the reference-marker convention. Validation failures may return NULL.
Recode reference-marker information using the package-specific convention
Description
Recode reference-marker information using the package-specific convention.
Usage
st.recode.ref.2(r1 = 0, r2 = 1, descending = FALSE,
data.set = "default")
Arguments
r1 |
Function argument used by this operation. |
r2 |
Function argument used by this operation. |
descending |
Function argument used by this operation. |
data.set |
Name of the SelectionTools data set. |
Value
On a successful run, an invisible implementation-level list returned by the final timing/information helper; the meaningful result is the side effect of recoding reference-marker information for two parents. Validation failures may return NULL.
Reset SelectionTools data and settings
Description
Reset SelectionTools data and settings.
Usage
st.reset()
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of resetting SelectionTools marker-data state.
st.restrict.marker.data()
Description
Restricts an existing marker data set by individual names, marker names, missingness, allele number, and expected heterozygosity.
Usage
st.restrict.marker.data(ind.list = NULL, ind.file = NULL, mar.list = NULL,
mar.file = NULL, NoAll.MAX = 999, MaMis.MAX = 1,
ExHet.MIN = 0, InMis.MAX = 1, data.set = "default")
Arguments
ind.list |
Optional character vector of individual names to retain. Unknown names are ignored. Do not supply together with |
ind.file |
Optional file containing individual names to retain. Relative names are read from the input directory. Do not supply together with |
mar.list |
Optional character vector of marker names to retain. Unknown names are ignored. Do not supply together with |
mar.file |
Optional file containing marker names to retain. Do not supply together with |
NoAll.MAX |
Non-negative maximum number of non-missing alleles allowed per marker. |
MaMis.MAX |
Finite maximum marker missingness. |
ExHet.MIN |
Finite minimum expected heterozygosity. |
InMis.MAX |
Finite maximum individual missingness. |
data.set |
Character string naming the marker data set to restrict. |
Details
The function first constructs individual and marker inclusion masks from the current marker statistics. Individual criteria and marker criteria are evaluated against the statistics that exist before the restriction starts. The filtering is therefore a single-pass operation, not an iterative filter-and-recalculate procedure. Calling the function repeatedly can be used when sequential recalculation between criteria is desired.
If an individual or marker selection list is supplied, names not found in the data set are silently ignored. Duplicate names in a selection list have no additional effect. The order of a selection list does not reorder the marker data: surviving individuals and markers remain in their original data-set order.
Marker filters use the number of non-missing alleles, marker missingness, and expected heterozygosity. Individual filtering uses individual missingness. All active criteria are combined. If every marker and every individual satisfies the criteria, the function is a true no-op and leaves all existing derived structures untouched.
A genuine restriction is transactional. A complete temporary marker matrix, phenotype vector, restricted map, and marker statistics are built before the current data set is replaced. If no individual or no marker would remain, or if construction fails, the existing data set is left unchanged.
If a map is present, it is restricted through the map access order and marker-to-data mapping, rather than by assuming that physical map storage has the same order as the marker matrix. This preserves valid map structures even when storage order and marker order differ. If markers are removed, marker-dependent model state is invalidated. If only individuals are removed, a complete and internally consistent marker-effect definition can be retained; incomplete or inconsistent effect state is not retained.
Value
On a successful run, an invisible implementation-level list returned by the final timing/information helper; the meaningful result is the side effect of restricting the marker data in the selected data set. Validation failures may return NULL.
Note
Restrictions are based on the marker statistics present at the start of each call. Repeating the function after statistics have been recalculated can therefore give a different result from applying several criteria simultaneously.
Return performance data from a SelectionTools data set
Description
Return performance data from a SelectionTools data set.
Usage
st.return.performance.data(out.filename = "y.vector", auxfiles = TRUE,
data.set ="default")
Arguments
out.filename |
Character string retained for compatibility and used as part of the name of an internal file in the R session temporary directory. The temporary file is removed before the function returns. |
auxfiles |
Logical. If |
data.set |
Name of the SelectionTools data set. |
Value
Invisibly returns a data.frame with columns i (individual identifier) and y (performance value) when the data can be returned; otherwise NULL.
Select individuals using phenotypic information
Description
Select individuals using phenotypic information.
Usage
st.select.phen(p, n=0, t=0, decreasing=TRUE, nsmall=1)
Arguments
p |
Function argument used by this operation. |
n |
Number of items or repetitions. |
t |
Generation or method-specific time parameter. |
decreasing |
Function argument used by this operation. |
nsmall |
Function argument used by this operation. |
Value
A data.frame containing the selected rows of p, sorted and filtered as requested, with an added character column descr combining individual identifier and phenotypic value.
Set haplotype-block information
Description
Set haplotype-block information.
Usage
st.set.hblocks(haplotype.list, hap.symbol = "s", out.filename = "blocks",
auxfiles = TRUE, data.set = "default")
Arguments
haplotype.list |
Function argument used by this operation. |
hap.symbol |
Function argument used by this operation. |
out.filename |
Character string retained for compatibility and used as part of the name of an internal file in the R session temporary directory. The temporary file is removed before the function returns. |
auxfiles |
Logical. If |
data.set |
Name of the SelectionTools data set. |
Value
Invisibly returns a data.frame with columns Chrom, Pos, Name, Class, and Markers describing the installed haplotype blocks when auxfiles = TRUE and the operation succeeds; otherwise NULL.
Set info level used by SelectionTools
Description
Set info level used by SelectionTools.
Usage
st.set.info.level(level)
Arguments
level |
Function argument used by this operation. |
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of setting the SelectionTools information level.
Set matrix ops used by SelectionTools
Description
Set matrix ops used by SelectionTools.
Usage
st.set.matrix.ops(select.s="mt", select.p="st")
Arguments
select.s |
Function argument used by this operation. |
select.p |
Function argument used by this operation. |
Value
NULL, returned invisibly. The function is called for its side effect of configuring matrix operations.
Set num threads used by SelectionTools
Description
Set num threads used by SelectionTools.
Usage
st.set.num.threads(active)
Arguments
active |
Function argument used by this operation. |
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of setting the number of computational threads.
Set openblas threads used by SelectionTools
Description
Set openblas threads used by SelectionTools.
Usage
st.set.openblas.threads(active)
Arguments
active |
Function argument used by this operation. |
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of setting the OpenBLAS thread count.
Transfer Marker Effects to the Simulation Backend
Description
Transfers estimated marker effects from a SelectionTools marker data set directly to the active simulation backend.
Usage
st.set.sim.ef(name="effect", data.set="default")
Arguments
name |
Character string naming the effect in the simulation backend. |
data.set |
Character string naming the SelectionTools marker data set that provides the estimated marker effects. |
Details
The function transfers marker effects directly between the genomic-selection and simulation data structures; no intermediate effect file is written.
With the default allele codes used by the genomic-selection routines,
gs.set.allele.codes() assigns the additive dosage codes 0, 1, and 2
to the genotypes with zero, one, and two copies of the effect allele,
respectively. For marker j with estimated effect u_j, the
marker contribution in the genomic-selection representation is therefore
0, u_j, or 2u_j. The default codes for genotypes
containing a missing allele are 0.5, 1, and 1.5.
The simulation backend stores an additive effect per allele. For each
transferred marker, the effect allele receives +u_j/2 and the
complementary allele receives -u_j/2. Before adjustment of the
general mean, the three diploid genotype contributions are consequently
-u_j, 0, and +u_j. This representation is equivalent to
changing the marker dosage from z_{ij} to
z^*_{ij}=z_{ij}-1.
For the genomic-selection representation, the estimated genetic value is
\hat{g}_i=\hat{\mu}_{012}+\sum_{j=1}^{m}z_{ij}\hat{u}_j.
Using z_{ij}=z^*_{ij}+1 gives
\hat{g}_i=
\left(\hat{\mu}_{012}+\sum_{j=1}^{m}\hat{u}_j\right)
+\sum_{j=1}^{m}z^*_{ij}\hat{u}_j.
Thus, the general mean stored for the simulation effect is
\hat{\mu}_{\mathrm{sim}}=
\hat{\mu}_{012}+\sum_{j=1}^{m}\hat{u}_j.
This is the coding-shift correction described by Stranden and Christensen (2011), Equation (2). In their general notation the change in the estimated general mean is the coding-shift vector transposed times the marker-effect vector. For the shift from 0/1/2 to -1/0/1, the coding-shift vector consists of ones and the correction is therefore the sum of marker effects.
This shift is distinct from allele-frequency centering. If a marker column is
centered as z^c_{ij}=z_{ij}-2p_j, the corresponding change of the
general mean is 2\sum_j p_j u_j. The simulation transfer uses
z^*_{ij}=z_{ij}-1, not allele-frequency centering. The two shifts are
the same for a marker only when p_j=0.5.
Only markers that are present in the active simulation map are transferred. Consequently, the sum in the intercept correction is taken over transferred markers only. Markers that are absent from the simulation map are ignored with a warning.
The exact equivalence described above applies to the default allele codes.
gs.set.allele.codes() permits power users to define other codes; an
arbitrary custom coding is not in general represented by the fixed
-u_j/2, +u_j/2 simulation allele effects.
Marker effects are optional. If data.set contains no estimated marker
effects, the function returns successfully without changing the simulation
effects. A partially inconsistent marker-effect state is treated as an error.
When marker effects are present, a diploid simulation genome and a linkage map must already be installed. Marker effects are connected to simulation loci by marker name. The data set supplying the effects therefore need not be the data set that originally supplied the simulation map.
If an effect with the requested name is already loaded, the operation
fails. After a new effect is installed successfully, stored population
genetic and phenotypic values are removed because their effect dimension is
no longer current.
Value
An invisible list returned by the underlying .C routine. Component r is the status value used by the simulation bridge; the function is primarily called for its side effect of transferring the requested data to the simulation backend.
References
Stranden I, Christensen OF (2011) Allele coding in genomic evaluation. Genetics Selection Evolution 43:25. See Equation (2).
See Also
gs.set.allele.codes, st.set.sim.gp, st.set.sim.mp,
st.set.sim.pp, st.set.simpop
Set sim gp used by SelectionTools
Description
Derives simulation genome parameters from a SelectionTools marker-data map and installs them in the simulation backend.
Usage
st.set.sim.gp(data.set="default")
Arguments
data.set |
Character string naming the SelectionTools marker data set that provides the linkage map. |
Details
The marker data set must contain a valid linkage map. Chromosome identifiers must define a contiguous sequence from chromosome 1 through the maximum chromosome number, and positions must be finite and non-negative. Chromosome lengths are derived from the mapped positions.
The SelectionTools marker matrix is diploid, so the bridge installs genome parameters compatible with two homologues. This is the genome-parameter stage of transferring a marker data set into the simulation representation. It does not by itself install map points or create a population.
Value
An invisible list returned by the underlying .C routine. Component r is the status value used by the simulation bridge; the function is primarily called for its side effect of transferring the requested data to the simulation backend.
Set sim mp used by SelectionTools
Description
Copies the mapped loci of a SelectionTools marker data set into the active simulation linkage map.
Usage
st.set.sim.mp(data.set="default")
Arguments
data.set |
Character string naming the SelectionTools marker data set that provides the map points. |
Details
Map points are transferred with chromosome, position, marker name, and class information. The transfer respects the SelectionTools distinction between physical map storage and map access order and preserves the source access relationship used by the marker matrix.
Simulation genome parameters should already be compatible with the marker-data map. This stage installs map points but does not create the simulation population.
Value
An invisible list returned by the underlying .C routine. Component r is the status value used by the simulation bridge; the function is primarily called for its side effect of transferring the requested data to the simulation backend.
Transfer a Marker Population to the Simulation Backend
Description
Creates a simulation population from the diploid marker matrix of a SelectionTools data set.
Usage
st.set.sim.pp(pop.name="default", data.set="default")
Arguments
pop.name |
Character string naming the simulation population to create. The name must contain 2 to 255 ASCII letters or digits. |
data.set |
Character string naming the SelectionTools marker data set supplying individuals and alleles. |
Details
A diploid simulation genome and linkage map must already be installed. The source marker data set does not need to contain the linkage map that was used to initialize the simulation. Instead, the active simulation loci are matched to rows of the source marker matrix by marker name.
Every marker in the active simulation linkage map must occur in the source marker matrix. The source data set may contain additional markers; these are ignored during population transfer. Consequently, populations from several compatible SelectionTools data sets can be added to one initialized simulation without reloading genome parameters or the linkage map.
For each imported individual, allele values are placed at the corresponding simulation loci. SelectionTools biological allele codes 1 through 254 are transferred unchanged, while the internal SelectionTools missing representation is converted to the simulation missing-allele indicator. After construction, the imported population is checked marker by marker against the source marker matrix. If verification fails, the newly created population is removed.
Compatibility of biological coding across independently supplied data sets is
the responsibility of the user. In particular, matching marker names are
assumed to identify the same loci and allele coding. The requested
pop.name must not already exist in the simulation backend.
Value
An invisible list returned by the underlying .C routine. Component r is the status value used by the simulation bridge; the function is primarily called for its side effect of transferring the requested data to the simulation backend.
See Also
st.set.sim.gp, st.set.sim.mp, st.set.sim.ef,
st.set.simpop
Initialize the Simulation Backend from a SelectionTools Data Set
Description
Initializes the simulation backend from a SelectionTools marker data set and creates a diploid simulation population.
Usage
st.set.simpop(pop.name, data.set="default", effect.name="effect")
Arguments
pop.name |
Character string naming the simulation population to create. The name must contain 2 to 255 ASCII letters or digits. |
data.set |
Character string naming the source SelectionTools marker data set. |
effect.name |
Character string naming transferred marker effects in the
simulation backend. The default is |
Details
The function first calls reset.all and then performs the bridge stages
in sequence: st.set.sim.gp, st.set.sim.mp,
st.set.sim.ef, and st.set.sim.pp. Processing stops immediately
if a stage reports failure.
The source data set must provide marker data and a valid linkage map for the
genome-parameter and map stages. Estimated marker effects are optional. If no
estimated marker effects are present, st.set.sim.ef succeeds without
installing an effect and population transfer continues normally.
If effects are present, the intercept and marker effects are transferred
directly to the simulation backend under effect.name; no intermediate
effect file is created.
The individual bridge functions can also be called separately. In particular,
a simulation may be initialized once with genome parameters, map, and optional
effects and then populated from additional compatible marker data sets using
repeated calls to st.set.sim.pp.
Value
Invisibly returns the list produced by the last completed simulation-bridge stage. Component r is the bridge status value. The main effect is initialization of the simulation backend from the SelectionTools data set.
See Also
st.set.sim.gp, st.set.sim.mp, st.set.sim.ef,
st.set.sim.pp
Create a simple graphical genotype plot
Description
Create a simple graphical genotype plot.
Usage
st.simple.ggt.plot(data.set="default", ...)
Arguments
data.set |
Name of the SelectionTools data set. |
... |
Additional arguments passed to the underlying operation. |
Value
NULL. The function is called for its side effect of creating one or more graphical-genotype PDF files.
Start or stop the SelectionTools timing helper
Description
Start or stop the SelectionTools timing helper.
Usage
st.start.timer(depth=1)
Arguments
depth |
Function argument used by this operation. |
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of starting the SelectionTools timing helper.
Start or stop the SelectionTools timing helper
Description
Start or stop the SelectionTools timing helper.
Usage
st.stop.timer(depth=1, info.level=0)
Arguments
depth |
Function argument used by this operation. |
info.level |
Function argument used by this operation. |
Value
An invisible implementation-level list returned by the final information-message call after elapsed time is calculated and reported.
Write linkage-map information for a SelectionTools data set
Description
Write linkage-map information for a SelectionTools data set.
Usage
st.write.map(mfilename, auxfiles = TRUE, data.set = "default")
Arguments
mfilename |
Required base name for the map file. The file is written with extension ‘.mmp’. |
auxfiles |
Logical. If |
data.set |
Name of the SelectionTools data set. |
Value
Invisibly returns a data.frame with columns Chrom, Pos, Name, and Class when auxfiles = TRUE and map output succeeds; otherwise NULL. The map file is also written as requested.
st.write.marker.data()
Description
Writes marker data, and where applicable map and genome information, in one of the supported external representations.
Usage
st.write.marker.data(format = "m", lfilename = NULL, mfilename = NULL,
nfilename = NULL, ifilename = "", f.ind = 0,
l.ind = 0, add.inum = ".1", auxfiles = FALSE,
data.set = "default")
Arguments
format |
Character string selecting |
lfilename |
Base name for list-format output. Required when |
mfilename |
Base name for matrix-format marker and map output. Required when |
nfilename |
Base name for NTSys/Plabsim marker, map, and genome-parameter output. Required when |
ifilename |
Optional file containing individual names to include in the output. |
f.ind |
First individual position to write when selecting by range. Positions are interpreted in the current data-set order. |
l.ind |
Last individual position to write when selecting by range. |
add.inum |
Character string appended to individual names in NTSys/Plabsim output. |
auxfiles |
Logical flag retained by the interface. The marker writer itself writes the files implied by |
data.set |
Character string naming the marker data set to write. |
Details
Format "l" writes list records consisting of individual, marker, and allele. A heterozygous genotype is written as two allele records. The internal append variant "a" appends list records to the list output file without writing a new header.
Format "m" writes a marker-major genotype matrix. When a linkage map is available, a corresponding map file is also written. Genotypes are written with an explicit slash between the two allele codes, which preserves multi-digit allele numbers unambiguously.
Format "n" writes the Plabsim/NTSys group of files: allele-incidence marker data, map data in Morgan units, and chromosome-length information. Individual names are adapted to the restrictions of this representation and add.inum is appended to the written name. Marker names containing a period are not suitable for this output because the period separates marker name and allele number. Name transformations that would make two output individuals indistinguishable are rejected.
The writer does not modify the in-memory marker data set. In particular, writing a data set without a linkage map uses the marker matrix directly and does not create a synthetic map or alter map dimensions. Unsupported format codes are rejected.
Individuals can be restricted either by a positional range or by a file of individual names. These output selections affect only the written files and do not alter the stored marker data.
Value
NULL, returned invisibly. The function is called for its side effect of writing marker-data files in the selected output format.
Fit a single-kernel mixed model using the SelectionTools native implementation
Description
Fit a single-kernel mixed model using the SelectionTools native implementation.
Usage
st_mixed(y, id = NULL, Z, Z_ids = NULL, pool1 = NULL, pool2 = NULL,
hybrid_id = NULL, include_sca = FALSE, ridge = 1e-8,
method = c("std", "vanraden", "sommer"),
coding = c("detect", "0 0.5 1", "-1 0 1", "0 1 2"), varcmp = NULL,
maxcyc = 100L, tol = 1e-10, calcVC = 1,
model_type = c("kernel", "marker"),
bootstrap = list(nObs=NULL, nMarkers=NULL, nRep=1, seed=NULL))
Arguments
y |
Response vector. |
id |
Identifier vector. |
Z |
Marker or design matrix. |
Z_ids |
Identifiers corresponding to rows of the marker matrix. |
pool1 |
Function argument used by this operation. |
pool2 |
Function argument used by this operation. |
hybrid_id |
Function argument used by this operation. |
include_sca |
Logical flag controlling inclusion of specific combining ability. |
ridge |
Ridge or numerical-stabilization parameter. |
method |
Method to use. |
coding |
Marker coding convention. |
varcmp |
Variance-component values. |
maxcyc |
Maximum number of fitting cycles. |
tol |
Convergence tolerance. |
calcVC |
Function argument used by this operation. |
model_type |
Function argument used by this operation. |
bootstrap |
Function argument used by this operation. |
Value
A named list returned by the native mixed-model implementation, or NULL if the native routine cannot construct a result. The component set depends on the selected mode.
For the single-population mode the list contains beta0 (intercept), varcmp (variance components), ghat (predicted genetic values), yhat (predicted response values), missing_ids (identifiers not matched to marker rows), diag (diagnostic counts and marker-coding information), and retval (compiled status code).
For hybrid modes the list contains beta0, varcmp, yhat_obs (predictions aligned to the observed-response vector), diag, retval, and ridge; a native implementation may additionally provide a debug component.
summarize.gvalue()
Description
Summarizes the genetic values of the individuals of a population with respect to defined effects
Usage
summarize.gvalue(pops, effectfile=NULL)
Arguments
pops |
Names separated by blanks of the populations to be summarized.) |
effectfile |
Name of the effect |
Details
If no effects are specified the weighted sum of all effects are used, otherwise only the single specified effects is used without using weights.
Value
A numeric matrix with one row per population and columns Obs, Mean, SDev, Min, Q10, Med, and Max, summarizing genetic values.
Note
See the lib00.example1 for an application of the command.
Swap population names
Description
Swap population names.
Usage
swap.population.name(NameP1, NameP2)
Arguments
NameP1 |
Name of the first parental population. |
NameP2 |
Name of the second parental population. |
Value
Returns the same invisible implementation-level list as population.swap.name(). The function is a compatibility wrapper and the meaningful result is the delegated population operation.
Enable or control SelectionTools console output
Description
Enable or control SelectionTools console output.
Usage
talk.to.me()
Value
Returns the same invisible implementation-level list as set.info.level(1); the main effect is enabling package information output.
Genetic Map for Tropical Maize
Description
Marker names, chromosome assignments, and map positions for tropical maize.
Usage
data("v-tropmaize-map")
Format
A data frame with 7140 rows and 3 variables:
nameCharacter marker name.
chromInteger chromosome number.
posNumeric marker position in the units used in the original map.
Source
Prepared from the original ‘ex-crossa.map’ text file.
Phenotypic Data for Tropical Maize
Description
Phenotypic observations for 264 tropical maize individuals.
Usage
data("v-tropmaize-phe")
Format
A data frame with 264 rows and 2 variables:
iiInteger individual identifier.
yyNumeric phenotypic value.
Source
Prepared from the original ‘ex-crossa.phe’ text file.
Genotype Data for Tropical Maize
Description
Compact, lossless representation of genotype data for 1135 markers and 264 tropical maize individuals.
Usage
data("v-tropmaize-pop")
Format
A list with 4 components:
markerCharacter vector of 1135 marker names.
individualCharacter vector of 264 individual identifiers.
genotype.levelsCharacter vector giving the genotype text represented by raw codes 0, 1, and 2, respectively:
"-1/-1","1/1", and"2/2".genotypeA 1135 by 264 raw matrix. Each element is a raw code in
as.raw(0:2)and indexesgenotype.levelsafter conversion to integer and addition of one. Rows correspond to markers and columns correspond to individuals.
Source
Prepared from the original ‘ex-crossa.pop’ text file.
Compact STvcf Genotype and Genetic-Map Data for Tropical Maize
Description
A compact STvcf representation of genotypes and genetic linkage positions for 264 tropical maize inbred lines. It contains markers occurring in both the tropical-maize marker matrix and the available genetic map.
Usage
data("v-tropmaize-vcf")
Format
A list of class "STvcf" with 1076 marker rows and 264 individuals:
CHROMInteger running chromosome codes 1 through 10.
POSDouble-precision genetic positions in cM retained from the available linkage map without intentional rounding.
IDCharacter vector containing 1076 unique marker IDs.
REFCharacter vector containing
"1"for every marker.ALTCharacter vector containing
"2"for every marker.individualCharacter vector
"1", ...,"264".allele1A 1076 by 264 raw matrix. Raw 0 denotes REF, raw 1 denotes ALT, and raw 255 denotes missing.
allele2A 1076 by 264 raw matrix with the same coding.
CHROM.levelsCharacter vector
"1", ...,"10".
Details
The original tropical-maize genotype object contains 1135 genotyped markers and the available linkage map contains 7140 entries. This STvcf object contains their 1076-marker intersection. Genotype-only markers lacking a map position and map-only loci lacking genotype data are not included.
POS contains the original double-precision cM positions from the distributed
linkage map. The positions are not rounded to integer cM. Exact ties are
permitted in STvcf; when installed with st.load.vcf.data(), ties are
separated internally by small deterministic offsets because the existing
SelectionTools map and simulation structures require distinct ordered locus
positions.
REF label "1" is VCF allele index 0 and ALT label "2" is VCF
allele index 1. Loading maps these indices to internal SelectionTools alleles
1 and 2 and maps raw 255 to the internal SelectionTools missing representation.
The marker and individual identifiers obey the strict STvcf/SelectionTools name convention used by the new converters: only ASCII letters and digits are permitted. The tropical-maize marker IDs are alphanumeric and the individual IDs are the digit strings 1 through 264.
Marker statistics and individual missing-data frequencies after loading are calculated from the 1076 markers in this object. Applying NoAll.MAX=2, MaMis.MAX=0.1, ExHet.MIN=0.1, and InMis.MAX=0.1 leaves 828 markers and 230 individuals, the intended tropical-maize analysis data set.
Source
Prepared from the tropical-maize genotype and linkage-map data distributed with SelectionTools.
Write version information in the package-specific format
Description
Write version information in the package-specific format.
Usage
write.version.2()
Value
An invisible list returned by the underlying .C routine. The list is an implementation-level result containing the arguments returned by the compiled routine; the function is primarily called for its side effect of requesting/printing package version information.