| Type: | Package |
| Title: | Interface to 'Biogeme' |
| Version: | 0.1.2 |
| Author: | Michel Bierlaire [aut, cre] |
| Maintainer: | Michel Bierlaire <michel.bierlaire@epfl.ch> |
| Description: | Uses the 'Python' implementation of 'Biogeme' as the numerical backend for specifying and estimating discrete-choice models in R. The default native requirement is 'biogeme==3.3.5'. |
| Depends: | R (≥ 4.3.0) |
| Imports: | reticulate (≥ 1.41.0), stats, utils |
| Suggests: | testthat (≥ 3.0.0), knitr, rmarkdown |
| SystemRequirements: | Python (>= 3.12), Biogeme (== 3.3.5) |
| License: | MIT + file LICENSE |
| Encoding: | UTF-8 |
| Config/testthat/edition: | 3 |
| VignetteBuilder: | knitr |
| URL: | https://github.com/michelbierlaire/rbiogeme |
| BugReports: | https://github.com/michelbierlaire/rbiogeme/issues |
| Date: | 2026-09-02 |
| Config/roxygen2/version: | 8.1.0 |
| NeedsCompilation: | no |
| Packaged: | 2026-09-19 09:20:39 UTC; bierlair |
| Repository: | CRAN |
| Date/Publication: | 2026-09-29 14:20:07 UTC |
Interface to Biogeme
Description
rbiogeme is a complete R-facing interface to the native Biogeme estimation engine. Data, parameters, expressions, models, estimation controls, and post-estimation operations are specified from R. The bridge compiles the complete expression tree once before native numerical work begins.
Details
Start a new session with biogeme_setup() to provision or verify the
runtime and receive corrective actions. Configure the runtime with
biogeme_config() before the first call that initializes Python when
selecting an existing interpreter. The default runtime is lazy. A typical workflow is
to create a numeric data frame, wrap it with biogeme_database(),
build symbolic expressions with variable() and
biogeme_beta(), construct a specialized or generic model, and call
estimate(). The returned fit supports the ordinary R methods
summary(), coef(), vcov(), logLik(), and
nobs().
Specialized constructors cover multinomial, nested, cross-nested, panel, Bayesian, MDCEV, hybrid-choice, catalog, assisted-specification, and sampled-alternative workflows. Generic models use a complete likelihood expression and can additionally hold a probability, named simulation expressions, weights, panel aggregation, draw metadata, subsets, and parameter overrides.
Expressions are symbolic. Arithmetic, comparisons, logical operators, transformations, probability functions, draws, integration nodes, and derivatives create neutral expression nodes; they do not evaluate a local R likelihood. Native Biogeme performs compilation, differentiation, integration, optimization, simulation, and reporting.
The expression operators +, -, *, /, and
^ build arithmetic nodes. Comparisons ==, !=,
<, <=, >, and >= build indicator nodes.
Logical conjunction, disjunction, and negation use \&, |, and
!. Use logzero() and safe_exp() when the native
numerically safe form is required.
Use estimate() for fresh estimation. Use
estimate_or_load() only when explicit YAML loading or recycling is
desired, and choose a fresh temporary output directory for equivalence tests.
Random-draw and sampling examples should fix the native draw design and seed
and document any remaining simulation noise.
Start with ?biogeme_setup or ?biogeme_check, then read
vignette("getting-started", package = "rbiogeme") and
vignette("modeling-workflows", package = "rbiogeme") for the R
syntax and the complete workflow. Advanced Bayesian, Monte Carlo, MDCEV,
catalog, hybrid-choice, and sampled-alternative examples are in
vignette("advanced-models", package = "rbiogeme").
The whole package was implemented by ChatGPT 5.6 Luna under the supervision of Michel Bierlaire.
Value
Depending on the function, a database, expression, model, diagnostics list, estimation result, or result-derived object.
See Also
biogeme_config, biogeme_database,
biogeme_model, logit_model, estimate,
simulate, summary.biogeme_fit, coef.biogeme_fit,
vcov.biogeme_fit, logLik.biogeme_fit
Select an expression from a numeric mapping
Description
Select an expression from a numeric mapping
Usage
Elem(mapping, index)
elem(mapping, index)
Arguments
mapping |
Named list whose names are integer keys. |
index |
Numeric or Biogeme expression used as the key. |
Value
A Biogeme expression.
Absolute value of a Biogeme expression
Description
Absolute value of a Biogeme expression
Usage
## S3 method for class 'biogeme_expression'
abs(x)
Arguments
x |
A scalar or Biogeme expression. |
Value
A Biogeme expression.
Run native Biogeme assisted specification
Description
The complete catalog search, quick-estimation objective evaluation, Pareto persistence, and final re-estimation are delegated to native Biogeme. The R interface accepts named objective and validity descriptors rather than exposing Python callback objects. The selected validity rule is created and executed inside the Python bridge between native estimates.
Usage
assisted_specification(
model,
objectives = "loglikelihood_dimension",
pareto_file_name = NULL,
model_name = "rbiogeme_assisted",
controls = list(),
force = TRUE,
control = NULL,
validity = NULL
)
run_assisted_specification(
model,
objectives = "loglikelihood_dimension",
pareto_file_name = NULL,
model_name = "rbiogeme_assisted",
controls = list(),
force = TRUE,
control = NULL,
validity = NULL
)
Arguments
model |
A biogeme_model containing catalog expressions. |
objectives |
Native objective preset. Supported values are "loglikelihood_dimension" and "aic_bic_dimension". |
pareto_file_name |
Explicit path of the native Pareto checkpoint file. It has no default because native checkpoints are persistent files. |
model_name |
Native Biogeme model name prefix. |
controls |
Named list of native Biogeme controls. |
force |
Whether to remove the named Pareto checkpoint and start fresh. |
control |
Optional biogeme_control() object; an alias for controls. |
validity |
Optional native validity-rule name. Currently supported is
|
Value
An object of class biogeme_assisted_fit containing native final results, the summary table, descriptions, and Pareto metadata.
Estimate a model with native Bayesian inference
Description
The complete R expression tree is compiled once, then native Biogeme builds and samples the PyMC model. The returned R object contains the native posterior summary and the paths of any generated NetCDF/YAML/HTML files; posterior draws are not represented as manually managed Python objects.
Usage
bayesian_estimate(
model,
model_name = "rbiogeme_model",
controls = list(),
starting_values = NULL,
control = NULL
)
Arguments
model |
A |
model_name |
Native Biogeme model name. |
controls |
Named native Biogeme controls. |
starting_values |
Optional named numeric vector of starting values. |
control |
Optional |
Details
NetCDF and YAML output are enabled by default for this operation. Additional
Bayesian controls such as bayesian_draws, warmup, chains, target_accept,
calculate_likelihood, calculate_waic, calculate_loo, and
mcmc_sampling_strategy are passed to native Biogeme through controls.
Because those default Bayesian files are persistent, supply an explicit
output_directory through biogeme_control() (or disable both output types
explicitly).
Value
An object of class biogeme_bayesian_fit containing the native
posterior summary and output paths.
Retrieve native posterior means by observation
Description
Loads the NetCDF result with Biogeme's public BayesianResults API and
delegates the observation-level posterior mean calculation to native Biogeme.
No posterior-draw or ArviZ object is exposed to R.
Usage
bayesian_posterior_mean_by_observation(bayesian_results, variable_name,
control = NULL)
Arguments
bayesian_results |
A |
variable_name |
Name of a stored posterior variable with one observation dimension in addition to chain and draw. |
control |
Optional native Biogeme control list. Likelihood, WAIC, and LOO calculations are disabled by default. |
Value
A data frame indexed by the native observation coordinate, with one
column named after variable_name.
Report variables stored in native Bayesian results
Description
Returns the serialized result of native Biogeme's
BayesianResultsSummary.report_stored_variables() operation. No PyMC or
ArviZ object is exposed to R.
Usage
bayesian_stored_variables(x)
Arguments
x |
A |
Value
A data frame describing each native stored variable by result group, variable name, dimensions, and shape.
Create a Biogeme parameter
Description
Create a Biogeme parameter
Usage
biogeme_beta(
name,
start = 0,
lower = NULL,
upper = NULL,
fixed = FALSE,
sigma_prior = 5,
prior = NULL
)
Arguments
name |
Parameter name. |
start |
Initial value. |
lower |
Optional lower bound. |
upper |
Optional upper bound. |
fixed |
Logical value indicating whether the parameter is fixed. |
sigma_prior |
Standard deviation of the native default normal prior. |
prior |
Optional declarative prior from |
Details
Parameter names are preserved exactly by the bridge and are therefore part
of the equivalence contract with native Biogeme. Bounds and fixed are
passed to the native parameter definition.
Value
A parameter expression.
Examples
b_cost <- biogeme_beta("b_cost", start = -1, upper = 0)
b_cost
Locate and import the package's Python bridge.
Description
This is intentionally an internal helper. Users interact with the high-level estimation functions; keeping the import in one place makes the Python boundary easy to test and keeps all numerical work on the Python side.
Usage
biogeme_bridge()
Create a native-compatible expression catalog
Description
Create a native-compatible expression catalog
Usage
biogeme_catalog(name, expressions, controller = NULL)
catalog(name, expressions, controller = NULL)
Arguments
name |
Catalog name. |
expressions |
Non-empty named list of alternative expressions. |
controller |
Optional |
Value
A catalog expression compiled to native Biogeme.
Create a neutral controller for one or more expression catalogs
Description
Catalog controllers select the same named specification in every catalog that shares the controller. The object remains an R specification until the complete expression is compiled to native Biogeme.
Usage
biogeme_catalog_controller(name, specification_names, selected_name = NULL)
catalog_controller(name, specification_names, selected_name = NULL)
Arguments
name |
Controller name. |
specification_names |
Ordered, unique catalog specification names. |
selected_name |
Optional specification selected before native compilation. |
Value
A neutral catalog controller.
Check whether rbiogeme is ready to run a model
Description
This is the runtime-only validation used by biogeme_setup() and is also
useful when the environment has already been configured. It initializes the
configured runtime, verifies the minimum R and Python versions, checks that
native Biogeme can be imported, and compares its version with the configured
requirement when that requirement pins an exact version. The check does not
construct or estimate a model.
Usage
biogeme_check(verbose = interactive())
Arguments
verbose |
If |
Value
An object of class biogeme_check with elements ready,
diagnostics, and issues. issues is a data frame with columns
check, status, message, and action.
Calculate native simulation confidence intervals
Description
Native Biogeme performs the repeated simulation and quantile calculation. R supplies ordinary parameter mappings and receives two ordinary data frames; no native simulation or result objects are exposed.
Usage
biogeme_confidence_intervals(
model,
beta_values,
expressions = NULL,
interval_size = 0.9,
control = NULL
)
Arguments
model |
A |
beta_values |
A non-empty list of named finite numeric parameter vectors, typically obtained from a fitted result's bootstrap field. |
expressions |
Optional named list of expressions. When omitted, the
model's |
interval_size |
Confidence interval size, strictly between zero and one. The native default is 0.9. |
control |
Optional |
Value
A list with left and right data frames and native metadata.
Configure the Python runtime used by rbiogeme
Description
Configuration must be performed before Python is initialized in the R
session. The default requirement is biogeme==3.3.5; use this function to
select a different compatible Biogeme requirement explicitly.
Usage
biogeme_config(python = NULL, biogeme_requirement = NULL, debug = NULL)
Arguments
python |
Optional Python executable. If |
biogeme_requirement |
Python requirement passed to
|
debug |
If |
Details
biogeme_config(python = ...) selects an existing Python interpreter; it
does not install Biogeme into that interpreter. If python is omitted,
reticulate can provision the configured requirement in its managed
environment.
Value
The current configuration, invisibly when changes are requested and visibly otherwise.
Define estimation and simulation controls
Description
Unspecified fields are omitted so native Biogeme defaults remain in force. Additional named arguments are passed through to the native bridge.
Usage
biogeme_control(
model_name = NULL,
output_directory = NULL,
seed = NULL,
numerically_safe = NULL,
second_derivatives = NULL,
second_derivatives_percentage = NULL,
optimization_algorithm = NULL,
number_of_draws = NULL,
draw_type = NULL,
draw_seed = NULL,
bootstrap_samples = NULL,
user_notes = NULL,
variance_covariance_type = NULL,
save_iterations = NULL,
generate_html = NULL,
generate_yaml = NULL,
validation_folds = NULL,
...
)
Arguments
model_name |
Optional native Biogeme model name. |
output_directory |
Optional explicit directory for native files. It is required whenever a control requests native HTML, YAML, NetCDF, pickle, or iteration output; no operation writes to the working directory by default. |
seed |
Optional random seed. |
numerically_safe |
Optional numerical-safety flag. |
second_derivatives |
Optional second-derivative policy. |
second_derivatives_percentage |
Optional native percentage of iterations using analytical second derivatives. |
optimization_algorithm |
Optional native algorithm name. |
number_of_draws |
Optional number of simulation draws. |
draw_type |
Optional default draw type metadata. |
draw_seed |
Optional draw seed (mapped to native |
bootstrap_samples |
Optional bootstrap sample count. |
user_notes |
Optional notes included in native reports/results. |
variance_covariance_type |
Optional result covariance selection, such
as |
save_iterations |
Optional iteration-file persistence flag. |
generate_html |
Optional HTML-report flag. |
generate_yaml |
Optional YAML-report flag. |
validation_folds |
Optional validation-fold count. |
... |
Additional native Biogeme controls. |
Details
Unspecified values are omitted. This lets native Biogeme apply its normal defaults. The control object can be stored in a model or passed to an estimation or simulation operation; it is not a Python object.
Value
A named list of controls.
Examples
biogeme_control(
seed = 1234,
generate_html = FALSE,
generate_yaml = FALSE,
save_iterations = FALSE
)
Create a Biogeme database
Description
The database remains an R object until a model is compiled. Data are copied and validated at construction time, so later modifications to the original data frame do not alter the model database.
Usage
biogeme_database(name, data)
Arguments
name |
Database name. |
data |
Numeric R data frame. |
Details
The input must contain only numeric, named, non-empty columns. The original row identifiers are retained separately from the data columns and are carried through native filtering and row-extraction operations.
Value
An object of class biogeme_database.
Examples
database <- biogeme_database(
"demo",
data.frame(choice = c(1, 2), income = c(10, 20))
)
biogeme_database_columns(database)
Return the columns available in a Biogeme database
Description
Return the columns available in a Biogeme database
Usage
biogeme_database_columns(database)
Arguments
database |
A |
Value
A character vector of column names.
Define a derived database variable using a complete Biogeme expression
Description
Define a derived database variable using a complete Biogeme expression
Usage
biogeme_database_define_variable(database, name, expression)
database_define_variable(database, name, expression)
define_variable(database, name, expression)
Arguments
database |
A |
name |
Name of the derived column. |
expression |
Complete Biogeme expression. |
Details
The expression is retained symbolically and is compiled by the Python bridge together with the model. It is not evaluated by an R callback during estimation or simulation.
Value
A new database specification.
Examples
database <- biogeme_database(
"demo",
data.frame(time = c(10, 12), cost = c(5, 6))
)
database <- biogeme_database_define_variable(
database, "time_cost", variable("time") / variable("cost")
)
biogeme_database_columns(database)
Extract observations from a Biogeme database
Description
This is the R equivalent of native Database.extract_rows(). Indices are
deliberately one-based, as in the rest of the R interface, and the
original row identifiers are retained in the returned database.
Usage
biogeme_database_extract_rows(database, indices)
database_extract_rows(database, indices)
Arguments
database |
A |
indices |
Unique positive one-based row indices, in the order to keep. |
Value
A new biogeme_database containing the selected observations.
Return the number of observations excluded by native database filters
Description
Return the number of observations excluded by native database filters
Usage
biogeme_database_filtered_row_count(database)
Arguments
database |
A |
Value
An integer, or NA until lazy filters are materialised.
Test whether a Biogeme database contains a column
Description
Test whether a Biogeme database contains a column
Usage
biogeme_database_has_column(database, column)
Arguments
database |
A |
column |
Column name. |
Value
A single logical value.
Return whether a database has a declared panel structure
Description
Return whether a database has a declared panel structure
Usage
biogeme_database_is_panel(database)
Arguments
database |
A |
Value
One logical value.
Materialize lazy database operations through native Biogeme
Description
Materialize lazy database operations through native Biogeme
Usage
biogeme_database_materialize(database)
Arguments
database |
A |
Value
A database with native derived columns and filters applied.
Return the number of rows currently represented by a database
Description
Return the number of rows currently represented by a database
Usage
biogeme_database_nrow(database)
Arguments
database |
A |
Value
An integer, or NA until lazy native filters are materialised.
Declare and validate a panel identifier
Description
Declare and validate a panel identifier
Usage
biogeme_database_panel(database, panel_id)
Arguments
database |
A |
panel_id |
Name of the panel identifier column. |
Details
The identifier must be numeric and finite. Observations belonging to one identifier must be contiguous; the constructor rejects interleaved panel rows rather than silently reordering them.
Value
A new panel database specification.
Examples
panel_database <- biogeme_panel_database(
"panel_demo",
data.frame(person = c(1, 1, 2, 2), choice = c(1, 2, 2, 1)),
panel_id = "person"
)
biogeme_database_is_panel(panel_database)
Remove observations satisfying a Biogeme logical expression
Description
Remove observations satisfying a Biogeme logical expression
Usage
biogeme_database_remove(database, condition)
database_remove(database, condition)
remove_observations(database, condition)
Arguments
database |
A |
condition |
Logical Biogeme expression; matching rows are removed. |
Details
condition is a symbolic indicator. Rows for which it is true are removed
when the database is materialized by native Biogeme.
Value
A new database specification.
Examples
database <- biogeme_database(
"demo",
data.frame(choice = c(1, 2, 1), income = c(10, 20, 30))
)
filtered <- biogeme_database_remove(database, variable("income") > 10)
biogeme_database_nrow(filtered)
Return the original row identifiers of a Biogeme database
Description
Return the original row identifiers of a Biogeme database
Usage
biogeme_database_row_ids(database)
Arguments
database |
A |
Value
A character vector of row identifiers.
Generate a native-compatible segmentation specification from a database
Description
Generate a native-compatible segmentation specification from a database
Usage
biogeme_database_segmentation(database, variable, mapping, reference = NULL)
database_generate_segmentation(database, variable, mapping, reference = NULL)
Arguments
database |
A |
variable |
A database variable name or expression. |
mapping |
Named mapping from integer values to segment names. |
reference |
Optional reference segment name. |
Value
A biogeme_segmentation specification.
Report the active R, Python, Biogeme, and numerical-library versions
Description
This function initializes the configured runtime. For a check that catches
initialization failures and returns an actionable status object, use
biogeme_check().
Usage
biogeme_diagnostics()
Value
A named list containing environment diagnostics.
See Also
biogeme_check(), biogeme_config()
Declare first-class draw metadata for a model
Description
Declare first-class draw metadata for a model
Usage
biogeme_draws(
name,
draw_type = "NORMAL",
number_of_draws = NULL,
seed = NULL,
matrix = NULL,
generator = NULL
)
Arguments
name |
Draw variable name. |
draw_type |
Native draw type. |
number_of_draws |
Optional draw count. |
seed |
Optional draw seed. |
matrix |
Optional user-supplied numeric draw matrix. |
generator |
Optional bridge generator name. Supported bridge-owned
generators include |
Value
A biogeme_draws object.
Return the diagnostic details attached to a Biogeme error
Description
Return the diagnostic details attached to a Biogeme error
Usage
biogeme_error_details(error)
Arguments
error |
A condition inheriting from |
Value
A named list with operation, suggestion, parent, and optional Python traceback details.
Create a Biogeme function node
Description
This constructor is intentionally generic so that new public Biogeme functions can be added without evaluating the expression in R.
Usage
biogeme_function(operator, args = list(), ...)
Arguments
operator |
Native Biogeme function/operator name. |
args |
A list of scalar values or Biogeme expressions. |
Value
A Biogeme expression.
Return native Biogeme general estimation statistics
Description
Return native Biogeme general estimation statistics
Usage
biogeme_general_statistics(fit)
Arguments
fit |
A |
Value
A named list of native Biogeme statistics.
Build synchronized generic/alternative-specific catalogs
Description
This is the neutral R counterpart of native Biogeme's
generic_alt_specific_catalogs(). Each returned element is a named list
of catalogs, one for every alternative. The catalogs share one native
controller with the exact specifications generic and altspec; when
segmentations are supplied, the segmentation catalogs are nested below
that controller in the same order as native Biogeme.
Usage
biogeme_generic_alt_specific_catalogs(
generic_name,
beta_parameters,
alternatives,
potential_segmentations = NULL,
maximum_number = 5L
)
generic_alt_specific_catalogs(
generic_name,
beta_parameters,
alternatives,
potential_segmentations = NULL,
maximum_number = 5L
)
Arguments
generic_name |
Shared name used for the segmentation and generic/alternative-specific controllers. |
beta_parameters |
Non-empty list of |
alternatives |
At least two unique alternative names. |
potential_segmentations |
Optional non-empty list of segmentation objects. |
maximum_number |
Maximum number of potential segmentations in one option. |
Value
A list of named lists of catalog expressions, one list per Beta.
Maximum of Biogeme expressions
Description
Maximum of Biogeme expressions
Usage
biogeme_max(...)
Arguments
... |
Scalar values or Biogeme expressions. |
Value
A Biogeme expression.
Construct an MDCEV model specification
Description
This is a declarative model object. The alternative utilities, shape parameters, and observed consumptions remain neutral R expression trees until the Python bridge sends the complete specification to the native Biogeme engine. The likelihood and forecasting algorithms are therefore provided by Biogeme itself.
Usage
biogeme_mdcev_model(
database,
model_type = c("gamma_profile", "generalized", "translated", "non_monotonic"),
baseline_utilities,
gamma_parameters,
alpha_parameters = NULL,
mu_utilities = NULL,
scale_parameter = NULL,
prices = NULL,
weights = NULL,
number_of_chosen_alternatives,
consumed_quantities,
subset = NULL,
control = NULL
)
mdcev_model(
database,
model_type = c("gamma_profile", "generalized", "translated", "non_monotonic"),
baseline_utilities,
gamma_parameters,
alpha_parameters = NULL,
mu_utilities = NULL,
scale_parameter = NULL,
prices = NULL,
weights = NULL,
number_of_chosen_alternatives,
consumed_quantities,
subset = NULL,
control = NULL
)
Arguments
database |
A |
model_type |
One of |
baseline_utilities |
Non-empty named list of baseline utility expressions, keyed by integer alternative code. |
gamma_parameters |
Named list of gamma expressions, keyed by the same
alternatives. One value may be |
alpha_parameters |
Named list of alpha expressions. Required for the generalized, translated, and non-monotonic variants. |
mu_utilities |
Named list of non-monotonic utility expressions. |
scale_parameter |
Optional scale expression. |
prices |
Optional named list of price expressions. Supported by the gamma-profile and generalized native classes. |
weights |
Optional observation-weight expression. |
number_of_chosen_alternatives |
Expression containing the number of goods chosen in each observation. |
consumed_quantities |
Named list of observed consumption expressions. |
subset |
Optional logical expression selecting observations to remove
when it is false, using the same semantics as |
control |
Optional |
Value
An object of class biogeme_mdcev_model.
Minimum of Biogeme expressions
Description
Minimum of Biogeme expressions
Usage
biogeme_min(...)
Arguments
... |
Scalar values or Biogeme expressions. |
Value
A Biogeme expression.
Create a generic Biogeme model
Description
The formula is a neutral R expression. It is compiled once into a native Biogeme expression graph before any estimation or simulation starts.
Usage
biogeme_model(
database,
formula = NULL,
weight = NULL,
probability = NULL,
simulations = NULL,
panel_trajectory = FALSE,
draws = NULL,
subset = NULL,
parameter_overrides = NULL,
control = NULL,
availability = NULL
)
Arguments
database |
A |
formula |
A log-likelihood expression, or a named list of formulas.
The names |
weight |
Optional observation-weight expression. |
probability |
Optional probability expression retained for simulation workflows. |
simulations |
Optional named list of expressions to simulate. |
panel_trajectory |
If |
draws |
Optional draw metadata object or list of draw metadata. |
subset |
Optional logical expression selecting observations to retain. |
parameter_overrides |
Optional named list of native parameter replacements, keyed by the original Beta names. |
control |
Optional |
availability |
Optional named list of availability expressions. This metadata is used by native post-estimation operations such as the null log-likelihood calculation for generic catalog models. |
Details
formula is the native log-likelihood expression. A named list may contain
log_like (or loglike) and additional expressions, but a generic model
must still provide a likelihood, probability, or simulation expressions.
simulations is a named list evaluated only when simulate() is called.
Value
An object of class biogeme_model.
Examples
database <- biogeme_database("demo", data.frame(choice = c(1, 2), x = c(1, 2)))
probability <- logit_probability(
utilities = list(`1` = 0, `2` = biogeme_beta("b") * variable("x")),
alternative = variable("choice")
)
model <- biogeme_model(
database,
formula = logzero(probability),
simulations = list(probability = probability)
)
model
Return the parameter definitions in a model
Description
Return the parameter definitions in a model
Usage
biogeme_model_parameters(model)
Arguments
model |
A Biogeme model. |
Value
A data frame of parameter definitions.
Return parameter names collected from the compiled native expression
Description
Unlike biogeme_model_parameters(), this operation also sees parameters
generated by native helpers such as segmentation and parameters in every
catalog branch. When configuration_id is supplied, native Biogeme first
resolves that configuration and the returned names describe the flattened
selected expression.
Usage
biogeme_native_parameter_names(
model,
configuration_id = NULL,
model_name = "rbiogeme_parameters",
controls = list()
)
Arguments
model |
A Biogeme model. |
configuration_id |
Optional exact native catalog configuration ID. |
model_name |
Temporary native model name used while compiling. |
controls |
Named native Biogeme controls. |
Value
A list with all, free, and fixed character vectors.
Construct a panel database in one call
Description
Construct a panel database in one call
Usage
biogeme_panel_database(name, data, panel_id)
Arguments
name |
Database name. |
data |
Numeric data frame. |
panel_id |
Panel identifier column. |
Value
A biogeme_database object.
Declare a native Bayesian prior
Description
Creates a data-only prior descriptor. The Python bridge compiles it into a native PyMC prior before estimation; R callbacks are not used by the sampler.
Usage
biogeme_prior(distribution = c("normal", "student_t"), sigma = 5, nu = 5)
Arguments
distribution |
Prior distribution: |
sigma |
Positive scale parameter. |
nu |
Positive degrees of freedom for a Student-t prior. |
Value
A declarative biogeme_prior object.
Initialize and return the Python interpreter used by rbiogeme
Description
Initialize and return the Python interpreter used by rbiogeme
Usage
biogeme_python()
Value
A reticulate Python configuration object. If no interpreter was selected, reticulate may provision the configured native requirement in its managed environment.
Define a native Biogeme sampling partition
Description
Stores the strata and sample sizes used by Biogeme's sampling-of-alternatives
API. The segments must cover full_set, and no sample size may exceed
the size of its segment.
Usage
biogeme_sampling_partition(segments, sample_sizes, full_set = NULL)
Arguments
segments |
A non-empty list of disjoint integer alternative-ID vectors. |
sample_sizes |
One positive sample size for each segment. |
full_set |
Optional integer vector containing the complete partition universe. |
Value
A declarative biogeme_sampling_partition object. Native Biogeme
performs the actual random sampling when the model is estimated.
See Also
Define a discrete parameter segmentation
Description
Define a discrete parameter segmentation
Usage
biogeme_segmentation(variable, mapping, reference = NULL)
Arguments
variable |
A data-variable name or expression. |
mapping |
Named mapping from integer values to segment names. |
reference |
Optional reference segment name. |
Value
A biogeme_segmentation specification.
Build synchronized catalogs for possible parameter segmentations
Description
This is the structural counterpart of native Biogeme's
segmentation_catalogs(). It only enumerates catalog names and builds
neutral expression nodes; native Biogeme still creates the segmented
parameters and evaluates the selected specification.
Usage
biogeme_segmentation_catalogs(
generic_name,
beta_parameters,
potential_segmentations,
maximum_number,
selected_name = NULL
)
segmentation_catalogs(
generic_name,
beta_parameters,
potential_segmentations,
maximum_number,
selected_name = NULL
)
Arguments
generic_name |
Shared controller name. |
beta_parameters |
Non-empty list of |
potential_segmentations |
Non-empty list of segmentation objects. |
maximum_number |
Maximum number of segmentations in one option. |
selected_name |
Optional segmentation specification selected before native compilation. |
Value
A list of synchronized catalog expressions, one per beta.
Automatically prepare the native Biogeme runtime
Description
This is the recommended first command for a new user. With no python
argument, reticulate provisions an isolated managed environment containing
the configured Biogeme requirement when the runtime is first initialized.
If python is supplied, that existing interpreter is selected instead and
must already contain the requested Biogeme package. In both cases the
function returns the same actionable status object as biogeme_check().
Usage
biogeme_setup(
python = NULL,
biogeme_requirement = NULL,
debug = NULL,
verbose = interactive()
)
Arguments
python |
Optional existing Python executable. If omitted, use the reticulate-managed environment when no user-managed environment has been selected. |
biogeme_requirement |
Python requirement to provision or verify.
Defaults to |
debug |
If |
verbose |
If |
Value
An object of class biogeme_check. Its ready element is TRUE
when the runtime can be used for model operations.
See Also
biogeme_check(), biogeme_config()
Box–Cox transformation
Description
Box–Cox transformation
Usage
boxcox(x, lambda)
Arguments
x |
Expression to transform. |
lambda |
Box–Cox exponent expression. |
Value
A Biogeme expression.
Return catalog configuration identifiers using native Biogeme
Description
Return catalog configuration identifiers using native Biogeme
Usage
catalog_configuration_ids(
model,
model_name = "rbiogeme_catalog",
controls = list(),
control = NULL
)
Arguments
model |
A |
model_name |
Native Biogeme model name. |
controls |
Named list of native Biogeme controls. |
control |
Optional |
Value
Character vector of exact native configuration identifiers.
Check analytical derivatives against native finite differences
Description
This delegates to the public BIOGEME.check_derivatives() operation. The
complete expression tree is compiled first; no R callback is evaluated by
the derivative checker.
Usage
check_derivatives(
model,
model_name = "rbiogeme_model",
controls = list(),
control = NULL,
verbose = FALSE
)
Arguments
model |
A |
model_name |
Native Biogeme model name. |
controls |
Named native Biogeme controls. |
control |
Optional |
verbose |
Whether native Biogeme should print the comparison. |
Value
A list containing the native function value, analytical and finite difference gradients and Hessians, and their error vectors.
Run native post-estimation Monte Carlo draw-stability diagnostics
Description
The model is compiled once and the supplied fixed estimate is passed to
native BIOGEME.check_monte_carlo_stability(). Native Biogeme evaluates the
objective and gradient at fresh draw designs and writes its YAML checkpoint
and Markdown report.
Usage
check_monte_carlo_stability(
model,
fit,
model_name = "rbiogeme_model",
controls = list(),
output_directory = NULL,
basename = NULL,
resume = TRUE,
control = NULL
)
Arguments
model |
A |
fit |
A |
model_name |
Native Biogeme model name. |
controls |
Named native Biogeme controls. |
output_directory |
Explicit directory for the native diagnostic files. It has no default: the operation refuses to write into the working directory when this argument is omitted. |
basename |
Optional native diagnostic filename prefix. |
resume |
Whether native Biogeme may resume a compatible checkpoint. |
control |
Optional |
Value
An object of class biogeme_monte_carlo_diagnostic containing the
native status, conclusion, recommendation, data, and output paths.
Collect parameter names from an expression
Description
Collect parameter names from an expression
Usage
collect_biogeme_parameters(expression)
Arguments
expression |
A Biogeme expression. |
Value
A character vector of parameter names.
Collect data-variable names from an expression
Description
Collect data-variable names from an expression
Usage
collect_biogeme_variables(expression)
Arguments
expression |
A Biogeme expression. |
Value
A character vector.
Count catalog specifications using native Biogeme
Description
Count catalog specifications using native Biogeme
Usage
count_number_of_specifications(
model,
model_name = "rbiogeme_catalog",
controls = list(),
control = NULL
)
Arguments
model |
A |
model_name |
Native Biogeme model name. |
controls |
Named list of native Biogeme controls. |
control |
Optional |
Value
The native number of catalog configurations, or NULL when native
Biogeme cannot determine a count.
Construct a native cross-nested-logit log-probability expression
Description
Construct a native cross-nested-logit log-probability expression
Usage
cross_nested_log_probability(
utilities,
availability,
nests,
alternative,
alternative_codes = NULL,
scale_parameter = NULL
)
Arguments
utilities |
Named list of utility expressions. |
availability |
Optional named list of availability expressions. |
nests |
A |
alternative |
Integer-valued alternative or a Biogeme choice expression. |
alternative_codes |
Optional named integer vector mapping utility names. |
scale_parameter |
Optional native CNL scale expression. |
Value
A Biogeme expression compiled to native biogeme.models.logcnl.
Calculate the native cross-nested-logit error-term correlation matrix
Description
Calculate the native cross-nested-logit error-term correlation matrix
Usage
cross_nested_logit_correlation(
model,
beta_values = NULL,
alternatives_names = NULL
)
Arguments
model |
A |
beta_values |
Optional named estimated or starting parameter values. |
alternatives_names |
Optional named character vector for row/column labels. |
Value
A numeric correlation matrix.
Define a cross-sectional cross-nested-logit model
Description
Define a cross-sectional cross-nested-logit model
Usage
cross_nested_logit_model(
database,
choice,
utilities,
nests,
availability = NULL,
alternative_codes = NULL,
scale_parameter = NULL,
control = NULL
)
Arguments
database |
A |
choice |
Name of the integer-valued choice column. |
utilities |
Named list of utility expressions. |
nests |
A |
availability |
Optional named list of availability expressions. |
alternative_codes |
Optional named integer vector mapping utility names. |
scale_parameter |
Optional native CNL scale expression. |
control |
Optional |
Value
An object of class biogeme_cross_nested_logit_model.
Define one cross-nested-logit nest
Description
Define one cross-nested-logit nest
Usage
cross_nested_nest(nest_parameter, allocation, name = NULL)
Arguments
nest_parameter |
Native nest parameter expression or numeric value. |
allocation |
Named mapping from alternative codes to allocation expressions in this nest. |
name |
Optional native nest name. |
Value
A biogeme_cross_nested_nest specification.
Define the complete cross-nested-logit nest structure
Description
Define the complete cross-nested-logit nest structure
Usage
cross_nested_nests(choice_set, nests, sparse = FALSE)
Arguments
choice_set |
All integer alternative codes. |
nests |
A non-empty list of |
sparse |
If |
Value
A biogeme_cross_nested_nests specification.
Construct a native cross-nested-logit probability expression
Description
Construct a native cross-nested-logit probability expression
Usage
cross_nested_probability(
utilities,
availability,
nests,
alternative,
alternative_codes = NULL
)
Arguments
utilities |
Named list of utility expressions. |
availability |
Optional named list of availability expressions. |
nests |
A |
alternative |
Integer-valued alternative or a Biogeme choice expression. |
alternative_codes |
Optional named integer vector mapping utility names. |
Value
A Biogeme expression compiled to native biogeme.models.cnl.
Summarize the structural sparsity of a CNL nest specification
Description
Summarize the structural sparsity of a CNL nest specification
Usage
cross_nested_sparsity_report(nests)
Arguments
nests |
A |
Value
A data frame with one row per nest and membership counts.
Define a sampled-alternative cross-variable
Description
Represents native Biogeme's CrossVariableTuple. The Python bridge
expands the expression after native choice-set sampling; R callbacks are not
used during estimation.
Usage
cross_variable(name, formula)
Arguments
name |
Name of the generated cross-variable. |
formula |
A neutral Biogeme expression involving individual and alternative variables. |
Value
A declarative biogeme_cross_variable object.
Symbolically differentiate an expression with respect to a named variable
Description
Symbolically differentiate an expression with respect to a named variable
Usage
derive(expression, name)
Derive(expression, name)
Arguments
expression |
Expression to differentiate. |
name |
Variable or parameter name. |
Value
A Biogeme expression.
Store a simulated individual-level parameter in Bayesian results
Description
Wraps a child expression in native Biogeme's
DistributedParameter node. The wrapper preserves the named variable in
Bayesian output and is compiled before native estimation; it is not evaluated
by R during sampling.
Usage
distributed_parameter(name, expression)
Arguments
name |
Name used for the stored native variable. |
expression |
Child Biogeme expression, usually a location parameter
plus a scale parameter multiplied by |
Details
This helper is a thin interface to native Biogeme and does not implement a sampling or likelihood engine in R.
Value
A symbolic Biogeme expression compiled to native
DistributedParameter.
Examples
beta <- biogeme_beta("b_time")
random_coefficient <- distributed_parameter(
"b_time_rnd",
beta + biogeme_beta("b_time_s", start = 1) * draw("b_time_eps", "NORMAL")
)
format(random_coefficient)
Create a named Biogeme draw node
Description
Create a named Biogeme draw node
Usage
draw(name, draw_type = "NORMAL")
Arguments
name |
Draw variable name. |
draw_type |
Biogeme draw type, such as |
Value
A Biogeme expression.
Estimate a Biogeme model
Description
Estimate a Biogeme model
Usage
estimate(
model,
model_name = "rbiogeme_model",
controls = list(),
starting_values = NULL,
run_bootstrap = FALSE,
yaml_file_name = NULL,
control = NULL
)
Arguments
model |
A |
model_name |
Output/model name used by Biogeme. |
controls |
Named list of Biogeme controls. |
starting_values |
Optional named numeric vector of starting values. |
run_bootstrap |
Whether to run bootstrap re-estimation. |
yaml_file_name |
Optional path for Biogeme's standard YAML output. |
control |
Optional |
Details
estimate() always requests a fresh native estimation. It does not load a
previous YAML or iteration file implicitly. Use estimate_or_load() when
explicit result loading or recycling is required.
Value
An object of class biogeme_fit.
Examples
## Not run:
fit <- estimate(
model,
model_name = "demo",
control = biogeme_control(
output_directory = tempfile("rbiogeme-estimate-"),
generate_html = FALSE,
generate_yaml = FALSE,
save_iterations = FALSE
)
)
summary(fit)
## End(Not run)
Estimate every specification in a native Biogeme catalog
Description
Catalog choices, synchronized controllers, estimation, result summaries, Pareto selection, and parameter comparison are all handled by native Biogeme. The returned object only contains ordinary R representations of those native results.
Usage
estimate_catalog(
model,
model_name = "rbiogeme_catalog",
controls = list(),
quick_estimate = FALSE,
run_bootstrap = FALSE,
force = TRUE,
control = NULL
)
Arguments
model |
A |
model_name |
Native Biogeme model name prefix. |
controls |
Named list of native Biogeme controls. |
quick_estimate |
Whether to use native quick estimation. |
run_bootstrap |
Whether to run native bootstrap re-estimation. |
force |
Whether to estimate afresh rather than recycle native files. |
control |
Optional |
Value
An object of class biogeme_catalog_fit containing native model
results, summaries, Pareto-optimal specifications, and LaTeX comparison.
Estimate one named native catalog configuration
Description
The configuration identifier is resolved by native
BIOGEME.from_configuration(). R only supplies the compiled expression
tree and receives ordinary serialized estimation results; no R expression
or callback is evaluated during estimation.
Usage
estimate_configuration(
model,
configuration_id,
model_name = "rbiogeme_configuration",
controls = list(),
starting_values = NULL,
run_bootstrap = FALSE,
yaml_file_name = NULL,
control = NULL
)
Arguments
model |
A |
configuration_id |
Exact native configuration identifier, for example
|
model_name |
Native Biogeme model name. |
controls |
Named list of native Biogeme controls. |
starting_values |
Optional named numeric vector of starting values. |
run_bootstrap |
Whether to run native bootstrap re-estimation. |
yaml_file_name |
Optional path for native YAML output. |
control |
Optional |
Value
An object of class biogeme_fit containing native results.
Estimate or explicitly load a standard Biogeme result
Description
force = FALSE permits loading the named YAML file; force = TRUE always
performs fresh estimation. The default estimate() path never recycles an
old result file.
Usage
estimate_or_load(
model,
yaml_file_name,
force = FALSE,
model_name = "rbiogeme_model",
controls = list(),
starting_values = NULL,
run_bootstrap = FALSE
)
Arguments
model |
A |
yaml_file_name |
YAML path used for loading or saving. |
force |
Whether to force fresh estimation. |
model_name |
Native Biogeme model name. |
controls |
Named native controls. |
starting_values |
Optional named starting values. |
run_bootstrap |
Whether to run bootstrap estimation. |
Value
A biogeme_fit object.
Estimate a native sampled-alternative model
Description
Regenerates the sampled choice sets with recycling disabled, compiles the sampled likelihood once in native Biogeme, and delegates estimation to the native engine. No Python source is generated by the R interface.
Usage
estimate_sampled_alternatives(
model,
model_name = "rbiogeme_sampled",
controls = list(),
starting_values = NULL,
run_bootstrap = FALSE,
control = NULL
)
Arguments
model |
A [ |
model_name |
Native Biogeme model name. |
controls |
Named native Biogeme controls. |
starting_values |
Optional named finite numeric vector. |
run_bootstrap |
Whether to run native bootstrap re-estimation. |
control |
Optional [ |
Value
A biogeme_fit object with native sampling-context and sampled-file
metadata attached.
See Also
Evaluate a standalone expression with native Biogeme
Description
The expression is compiled once by the Python bridge. The initial result
contains the native function value, gradient, Hessian, and BHHH matrix. If
points is supplied, the same compiled native callable evaluates each row
with the requested free-parameter values. R does not evaluate the
expression locally and no Python object is returned.
Usage
evaluate_biogeme_expression(
expression,
beta = NULL,
points = NULL,
numerically_safe = FALSE,
use_jit = TRUE
)
Arguments
expression |
A Biogeme expression. |
beta |
Optional named numeric vector used for the initial evaluation. When omitted, native Beta starting values are used. |
points |
Optional data frame or numeric matrix. Its columns are free parameter names and its rows are the parameter vectors for repeated native evaluations. |
numerically_safe |
Whether to request native numerically safe formulas. |
use_jit |
Whether to use native JAX just-in-time compilation. |
Value
An object of class biogeme_expression_evaluation with initial,
free_beta_names, and evaluations components.
Evaluate one expression with native row-wise or aggregated JAX calculation
Description
This operation delegates to Biogeme's public get_value_c calculator. The
expression is compiled before the native evaluator runs; R receives only a
numeric vector or scalar.
Usage
evaluate_biogeme_expression_c(
model,
expression,
beta,
aggregation = FALSE,
number_of_draws = 1000L,
numerically_safe = FALSE,
use_jit = TRUE
)
Arguments
model |
A |
expression |
A Biogeme expression to evaluate. |
beta |
A |
aggregation |
Whether to return one native aggregated scalar instead of one value per observation. |
number_of_draws |
Positive number of native Monte Carlo draws. |
numerically_safe |
Whether to request native numerically safe formulas. |
use_jit |
Whether to use native JAX just-in-time compilation. |
Value
A numeric vector, or one numeric scalar when aggregation = TRUE.
Integrate an expression over a standard normal random variable
Description
Integrate an expression over a standard normal random variable
Usage
integrate_normal(expression, name, number_of_quadrature_points = 30L)
Arguments
expression |
Integrand expression. |
name |
Random variable name. |
number_of_quadrature_points |
Number of quadrature points. |
Value
A Biogeme expression.
Define one coefficient-variable term for a linear utility
Description
Define one coefficient-variable term for a linear utility
Usage
linear_term(beta, x)
LinearTermTuple(beta, x)
Arguments
beta |
A |
x |
A |
Value
A linear-utility term.
Construct a native linear utility expression
Description
Construct a native linear utility expression
Usage
linear_utility(terms)
LinearUtility(terms)
Arguments
terms |
A non-empty list of |
Value
A Biogeme expression compiled to native LinearUtility.
Construct a native logit log-probability expression
Description
This is the log-probability counterpart of logit_probability(). It is
compiled directly to Biogeme's public models.loglogit constructor, which
preserves the native numerically stable likelihood used by examples such
as Swissmetro b20.
Usage
logit_log_probability(
utilities,
availability = NULL,
alternative,
alternative_codes = NULL
)
Arguments
utilities |
Named list of utility expressions keyed by alternatives. |
availability |
Optional named list of availability expressions. |
alternative |
Integer code or expression for the observed alternative. |
alternative_codes |
Optional named integer codes for the utilities. |
Value
A Biogeme logit log-probability expression.
Define a cross-sectional multinomial logit model
Description
Define a cross-sectional multinomial logit model
Usage
logit_model(
database,
choice,
utilities,
availability = NULL,
alternative_codes = NULL,
weight = NULL
)
Arguments
database |
A |
choice |
Name of the integer-valued choice column. |
utilities |
Named list of utility expressions, one per alternative. |
availability |
Optional named list of availability expressions. |
alternative_codes |
Optional named integer vector mapping utility names to values in the choice column. If omitted, utility names must themselves be integer codes. |
weight |
Optional observation-weight expression. |
Details
The utility list is named by alternative. If the names are not integer
codes, pass alternative_codes as a named integer vector. Availability and
weights may be constants or symbolic expressions.
Value
An object of class biogeme_logit_model.
Examples
database <- biogeme_database(
"demo",
data.frame(choice = c(1, 2), x = c(1, 2))
)
b <- biogeme_beta("b")
model <- logit_model(
database,
choice = "choice",
utilities = list(`1` = 0, `2` = b * variable("x"))
)
model
Construct a native logit probability expression
Description
Construct a native logit probability expression
Usage
logit_probability(
utilities,
availability = NULL,
alternative,
alternative_codes = NULL
)
logit(utilities, availability = NULL, alternative, alternative_codes = NULL)
Arguments
utilities |
Named list of utility expressions, one per alternative. |
availability |
Optional named list of availability expressions. |
alternative |
Integer-valued alternative whose probability is returned, or a Biogeme expression selecting the alternative for each observation. |
alternative_codes |
Optional named integer vector when utility names are not themselves integer alternative codes. |
Value
A probability expression compiled to native biogeme.models.logit.
A numerically safe logarithm that returns zero at zero
Description
A numerically safe logarithm that returns zero at zero
Usage
logzero(x)
safe_log(x)
Arguments
x |
A scalar or Biogeme expression. |
Value
A Biogeme expression.
Estimate an MDCEV model with native Biogeme
Description
The MDCEV likelihood is generated by the native Biogeme MDCEV class after
compilation. This convenience wrapper has the same fresh-estimation
semantics as estimate().
Usage
mdcev_estimate(
model,
model_name = "rbiogeme_mdcev",
controls = list(),
starting_values = NULL,
run_bootstrap = FALSE,
yaml_file_name = NULL,
control = NULL
)
Arguments
model |
A |
model_name |
Native Biogeme model name. |
controls |
Named native Biogeme controls. |
starting_values |
Optional named numeric vector of starting values. |
run_bootstrap |
Whether to run native bootstrap re-estimation. |
yaml_file_name |
Optional path for native YAML output. |
control |
Optional |
Value
A biogeme_fit object.
Forecast MDCEV consumption using native Biogeme algorithms
Description
Forecast MDCEV consumption using native Biogeme algorithms
Usage
mdcev_forecast(
model,
fit,
database = NULL,
total_budget,
epsilons = NULL,
number_of_draws = NULL,
seed = NULL,
brute_force = FALSE,
tolerance_dual = 1e-10,
tolerance_budget = 1e-10
)
Arguments
model |
A |
fit |
A fresh |
database |
Optional database used only for forecasting. This permits forecasting a native row subset without changing the estimation fit. |
total_budget |
Positive total budget. |
epsilons |
Optional list of Gumbel-draw matrices. If omitted, native
draws are generated using |
number_of_draws |
Number of draws when |
seed |
Optional native NumPy seed when draws are generated. |
brute_force |
Whether to use native brute-force optimization. |
tolerance_dual |
Native dual-variable tolerance. |
tolerance_budget |
Native budget-constraint tolerance. |
Value
A biogeme_mdcev_forecast object containing R data frames.
Return native pandas description tables for an MDCEV forecast
Description
Return native pandas description tables for an MDCEV forecast
Usage
mdcev_forecast_describe(forecast)
Arguments
forecast |
A |
Value
A list of ordinary R data frames corresponding to native
pandas DataFrame.describe() output.
Generate native MDCEV error-term draws
Description
Generate native MDCEV error-term draws
Usage
mdcev_generate_epsilons(
model,
number_of_observations,
number_of_draws,
seed = NULL
)
Arguments
model |
A |
number_of_observations |
Number of database observations. |
number_of_draws |
Number of Gumbel draws per observation. |
seed |
Optional native NumPy seed. |
Value
A list of numeric matrices, one matrix per observation.
Return native Biogeme's estimated-parameter table for an MDCEV fit
Description
Return native Biogeme's estimated-parameter table for an MDCEV fit
Usage
mdcev_parameter_table(fit, variance_covariance_type = NULL)
Arguments
fit |
A |
variance_covariance_type |
Optional native covariance type, such as
|
Value
A named list of ordinary R data frames returned by native Biogeme.
Return native Biogeme's compact MDCEV estimation summary
Description
Return native Biogeme's compact MDCEV estimation summary
Usage
mdcev_short_summary(fit)
Arguments
fit |
A |
Value
The native Biogeme summary string.
Validate the two native MDCEV forecasting algorithms
Description
Validate the two native MDCEV forecasting algorithms
Usage
mdcev_validate_forecast(
model,
fit,
database = NULL,
total_budget,
epsilons = NULL,
number_of_draws = NULL,
seed = NULL,
tolerance_dual = 1e-13,
tolerance_budget = 1e-13
)
Arguments
model |
A |
fit |
A fresh |
database |
Optional database used only for forecast validation. |
total_budget |
Positive total budget. |
epsilons |
Optional list of Gumbel-draw matrices. |
number_of_draws |
Number of native draws when |
seed |
Optional native NumPy seed when draws are generated. |
tolerance_dual |
Native dual-variable tolerance. |
tolerance_budget |
Native budget-constraint tolerance. |
Value
TRUE when native validation completes without an exception.
Monte Carlo average of an expression
Description
Monte Carlo average of an expression
Usage
monte_carlo(expression)
Arguments
expression |
Expression to integrate. |
Value
A Biogeme expression.
Construct a native nested-logit log probability with endogenous-sampling correction
Description
Construct a native nested-logit log probability with endogenous-sampling correction
Usage
nested_endogenous_sampling_log_probability(
utilities,
availability,
nests,
correction,
alternative,
alternative_codes = NULL
)
Arguments
utilities |
Named list of utility expressions. |
availability |
Optional named list of availability expressions. |
nests |
A |
correction |
Named list of alternative-specific log correction terms. |
alternative |
Integer-valued alternative or a Biogeme choice expression. |
alternative_codes |
Optional named integer vector mapping utility names. |
Value
A Biogeme expression compiled through native get_mev_for_nested
and logmev_endogenous_sampling.
Construct a native nested-logit log-probability expression
Description
Construct a native nested-logit log-probability expression
Usage
nested_log_probability(
utilities,
availability,
nests,
alternative,
alternative_codes = NULL,
scale_parameter = NULL
)
Arguments
utilities |
Named list of utility expressions. |
availability |
Optional named list of availability expressions. |
nests |
A |
alternative |
Integer-valued alternative or a Biogeme choice expression. |
alternative_codes |
Optional named integer vector mapping utility names. |
scale_parameter |
Optional native scale expression for bottom normalization. |
Value
A Biogeme expression compiled to native biogeme.models.lognested.
Calculate the native nested-logit error-term correlation matrix
Description
Calculate the native nested-logit error-term correlation matrix
Usage
nested_logit_correlation(
model,
beta_values = NULL,
alternatives_names = NULL,
mu = 1
)
Arguments
model |
A |
beta_values |
Optional named estimated or starting parameter values. |
alternatives_names |
Optional named character vector for row/column labels. |
mu |
Overall scale parameter passed to native Biogeme. |
Value
A numeric correlation matrix.
Define a cross-sectional nested-logit model
Description
Define a cross-sectional nested-logit model
Usage
nested_logit_model(
database,
choice,
utilities,
nests,
availability = NULL,
alternative_codes = NULL,
scale_parameter = NULL,
control = NULL
)
Arguments
database |
A |
choice |
Name of the integer-valued choice column. |
utilities |
Named list of utility expressions. |
nests |
A |
availability |
Optional named list of availability expressions. |
alternative_codes |
Optional named integer vector mapping utility names. |
scale_parameter |
Optional native scale expression for bottom normalization. |
control |
Optional |
Value
An object of class biogeme_nested_logit_model.
Define one non-trivial nested-logit nest
Description
Define one non-trivial nested-logit nest
Usage
nested_nest(nest_parameter, alternatives, name = NULL)
Arguments
nest_parameter |
Native nest parameter expression or numeric value. |
alternatives |
Integer alternative codes contained in the nest. |
name |
Optional native nest name. |
Value
A biogeme_nested_nest specification.
Define the complete nested-logit nest structure
Description
Alternatives not listed in a non-trivial nest are treated by native Biogeme as trivial nests containing one alternative.
Usage
nested_nests(choice_set, nests)
Arguments
choice_set |
All integer alternative codes. |
nests |
A non-empty list of |
Value
A biogeme_nested_nests specification.
Construct a native nested-logit probability expression
Description
Construct a native nested-logit probability expression
Usage
nested_probability(
utilities,
availability,
nests,
alternative,
alternative_codes = NULL
)
Arguments
utilities |
Named list of utility expressions. |
availability |
Optional named list of availability expressions. |
nests |
A |
alternative |
Integer-valued alternative or a Biogeme choice expression. |
alternative_codes |
Optional named integer vector mapping utility names. |
Value
A Biogeme expression compiled to native biogeme.models.nested.
Create a Biogeme expression node
Description
This is an internal constructor. Expressions are represented in R until a complete model is compiled by the Python bridge.
Usage
new_biogeme_expression(kind, ...)
Arguments
kind |
Expression-node kind. |
... |
Node attributes. |
Value
An object of class biogeme_expression.
Normal cumulative distribution function
Description
Normal cumulative distribution function
Usage
normal_cdf(x)
Arguments
x |
A scalar or Biogeme expression. |
Value
A Biogeme expression.
Normal probability density function
Description
Normal probability density function
Usage
normal_pdf(x)
Arguments
x |
A scalar or Biogeme expression. |
Value
A Biogeme expression.
Construct an ordered-logit log-likelihood expression
Description
Constructs a neutral expression node. The complete node is compiled once to Biogeme's public ordered-logit expression before estimation or simulation.
Usage
ordered_logit_log_probability(eta, cutpoints, alternative,
categories = NULL, neutral_labels = numeric(),
enforce_order = TRUE, eps = 1e-12)
Arguments
eta |
Latent-index expression. |
cutpoints |
Non-empty list of cutpoint expressions. |
alternative |
Observed response expression. |
categories |
Optional numeric response labels. |
neutral_labels |
Optional numeric labels with unit contribution. |
enforce_order |
Whether native Biogeme enforces ordered cutpoints. |
eps |
Positive numerical lower bound used by native Biogeme. |
Value
A Biogeme expression compiled to native OrderedLogLogit.
Construct an ordered-probit log-likelihood expression
Description
Constructs a neutral expression node. The complete node is compiled once to Biogeme's public ordered-probit expression before estimation or simulation.
Usage
ordered_probit_log_probability(eta, cutpoints, alternative,
categories = NULL, neutral_labels = numeric(),
enforce_order = TRUE, eps = 1e-12)
Arguments
eta |
Latent-index expression. |
cutpoints |
Non-empty list of cutpoint expressions. |
alternative |
Observed response expression. |
categories |
Optional numeric response labels. |
neutral_labels |
Optional numeric labels with unit contribution. |
enforce_order |
Whether native Biogeme enforces ordered cutpoints. |
eps |
Positive numerical lower bound used by native Biogeme. |
Value
A Biogeme expression compiled to native OrderedLogProbit.
Construct a native ordered-response log-likelihood expression
Description
The observed response is matched to categories; the intervals are
defined by the ordered cutpoints. This internal helper selects the
native logistic or normal cumulative distribution. No probability is
evaluated in R.
Usage
ordered_response_log_probability(
distribution = c("logit", "probit"),
eta,
cutpoints,
alternative,
categories = NULL,
neutral_labels = numeric(),
enforce_order = TRUE,
eps = 1e-12
)
Arguments
distribution |
Native ordered-response distribution. |
eta |
Latent-index expression. |
cutpoints |
Non-empty list of cutpoint expressions. |
alternative |
Observed response expression. |
categories |
Optional ordered numeric labels. If omitted, native
Biogeme uses |
neutral_labels |
Optional numeric labels whose contribution is one. |
enforce_order |
Whether native Biogeme should enforce ordered cutpoints during evaluation. |
eps |
Positive numerical lower bound used by native Biogeme. |
Value
A Biogeme expression compiled to native Biogeme.
Aggregate an observation-level likelihood over a panel trajectory
Description
Aggregate an observation-level likelihood over a panel trajectory
Usage
panel_likelihood_trajectory(expression)
Arguments
expression |
Observation-level likelihood or probability expression. |
Value
A Biogeme expression.
Re-estimate the Pareto-optimal models saved by native Biogeme
Description
The Pareto file is read by native ParetoPostProcessing. Re-estimation,
summary compilation, model descriptions, and optional Pareto plotting all
remain native operations; the returned object contains ordinary R values.
Usage
pareto_post_processing(
model,
pareto_file_name,
model_name = "rbiogeme_pareto",
controls = list(),
recycle = FALSE,
plot_file_name = NULL,
objective_x = 0L,
objective_y = 1L,
label_x = NULL,
label_y = NULL,
control = NULL
)
Arguments
model |
A |
pareto_file_name |
Path to an existing native Pareto file. |
model_name |
Native Biogeme model-name prefix for re-estimation. |
controls |
Named list of native Biogeme controls. |
recycle |
Whether native Biogeme may recycle complete estimation files.
The default is |
plot_file_name |
Optional path where native Pareto plot output is saved. |
objective_x |
Zero-based native Pareto objective index for the plot. |
objective_y |
Zero-based native Pareto objective index for the plot. |
label_x |
Optional native plot x-axis label. |
label_y |
Optional native plot y-axis label. |
control |
Optional |
Value
An object of class biogeme_pareto_fit containing native results,
the compiled summary, descriptions, Pareto statistics, and plot path.
Piecewise-linear expression using native Biogeme naming rules
Description
Piecewise-linear expression using native Biogeme naming rules
Usage
piecewise(
expression,
thresholds,
betas = NULL,
transform = c("formula", "variable")
)
Arguments
expression |
Variable expression. |
thresholds |
Piecewise thresholds; only endpoints may be NULL. |
betas |
Optional list of parameter expressions. |
transform |
Whether to return the full formula or the transformed variable form. |
Value
A Biogeme expression.
Predict native probabilities or simulation expressions
Description
Predict native probabilities or simulation expressions
Usage
## S3 method for class 'biogeme_fit'
predict(object, newdata = NULL, expressions = NULL, control = NULL, ...)
Arguments
object |
A |
newdata |
Optional numeric data frame or |
expressions |
Optional named list of complete symbolic expressions. For logit, nested-logit, and cross-nested-logit fits, omission produces one native probability column per alternative. For a generic model, the model's probability or simulation expressions are used. |
control |
Optional |
... |
Reserved for future prediction options; unsupported arguments are rejected explicitly. |
Details
predict() is a convenience wrapper around simulate(). It does not
calculate probabilities in R and does not refit the model. Use
simulate() directly when several named quantities, posterior draws, or a
custom simulation workflow are needed.
Value
A data frame containing values calculated by native Biogeme.
Examples
database <- biogeme_database(
"demo",
data.frame(choice = c(1, 2), x = c(1, 2))
)
model <- logit_model(
database,
choice = "choice",
utilities = list(`1` = 0, `2` = biogeme_beta("b") * variable("x"))
)
# After fitting: predict(fit) or predict(fit, newdata = data.frame(...))
Profile native JAX formula evaluation
Description
The complete expression tree is compiled once, then native Biogeme's
CompiledFormulaEvaluator is used for each requested evaluation mode.
Each mode is evaluated twice, matching the native profiling example's first
call and steady-state measurements. Only ordinary environment and timing
summaries cross back to R.
Usage
profile_jax(
model,
beta_values = NULL,
cases = list(
list(
label = "Function only", gradient = FALSE, hessian = FALSE, bhhh = FALSE
),
list(
label = "Function + gradient", gradient = TRUE, hessian = FALSE, bhhh = FALSE
),
list(
label = "Function + gradient + Hessian",
gradient = TRUE, hessian = TRUE, bhhh = FALSE
),
list(
label = "Function + gradient + BHHH",
gradient = TRUE, hessian = FALSE, bhhh = TRUE
)
),
model_name = "rbiogeme_model",
controls = list(),
numerically_safe = FALSE,
control = NULL
)
Arguments
model |
A |
beta_values |
Named finite numeric vector of parameter values. If |
cases |
A non-empty list of named case lists. Each case has a non-empty |
model_name |
Native Biogeme model name. |
controls |
Named native Biogeme controls. |
numerically_safe |
Whether native formula evaluation uses numerical-safety transformations. |
control |
Optional |
Details
Timing values depend on the active Python, JAX, hardware, and thread configuration. The operation is a profiling diagnostic, not an estimation operation.
Value
An object of class biogeme_jax_profile containing the native JAX environment and serialized profile data for each case.
See Also
biogeme_control
Estimate a Biogeme model using the native quick-estimation operation
Description
quick_estimate() delegates to native BIOGEME.quick_estimate(). It
returns parameter values and the statistics that native Biogeme makes
available without the ordinary post-estimation derivative calculations.
Use save_results() when a standard YAML result is required.
Usage
quick_estimate(
model,
model_name = "rbiogeme_model",
controls = list(),
control = NULL
)
Arguments
model |
A |
model_name |
Output/model name used by Biogeme. |
controls |
Named list of Biogeme controls. |
control |
Optional |
Value
An object of class biogeme_fit.
Create a random variable for native numerical integration
Description
Create a random variable for native numerical integration
Usage
random_variable(name)
Arguments
name |
Random variable name. |
Value
A Biogeme expression.
Read standard Biogeme YAML results
Description
Read standard Biogeme YAML results
Usage
read_results(filename)
Arguments
filename |
YAML file generated by Biogeme. |
Value
A biogeme_fit object.
Numerically safe exponential
Description
Numerically safe exponential
Usage
safe_exp(x)
Arguments
x |
A scalar or Biogeme expression. |
Value
A Biogeme expression.
Define a sampled-alternative Biogeme model
Description
Defines the protocol for a native sampling-of-alternatives model. Sampling, cross-variable expansion, likelihood construction, differentiation, and optimization are performed by native Biogeme through the Python bridge. Sampling is regenerated with recycling disabled by default.
Usage
sampled_alternatives_model(
alternatives,
individuals,
choice_column,
id_column,
utility,
partition,
biogeme_file_name,
model_type = c("logit", "nested", "cnl"),
combined_variables = list(),
mev_partition = NULL,
mev_sample_sizes = NULL,
nests = NULL,
control = NULL
)
Arguments
alternatives |
Numeric data frame with one row per alternative. |
individuals |
Numeric data frame with one row per decision maker. |
choice_column |
Choice column in |
id_column |
Unique alternative-ID column in |
utility |
Complete neutral Biogeme utility expression. |
partition |
Main [ |
biogeme_file_name |
Path for the native sampled-data file. |
model_type |
One of |
combined_variables |
Optional list of [ |
mev_partition |
Optional partition used for MEV terms. |
mev_sample_sizes |
Optional MEV sample sizes. |
nests |
Nested or cross-nested nest object for the selected model type. |
control |
Optional [ |
Value
A biogeme_sampled_alternatives_model object containing ordinary R
data and a neutral expression specification.
See Also
estimate_sampled_alternatives()
Generate balanced sampling segment sizes
Description
Creates sampling protocol metadata equivalent to native
generate_segment_size(); native Biogeme still performs validation and
sampling.
Usage
sampling_segment_sizes(sample_size, number_of_segments)
Arguments
sample_size |
Total number of alternatives to sample. |
number_of_segments |
Number of sampling segments. |
Value
An integer vector whose entries differ by at most one. Any remainder is assigned to the first segments, matching native Biogeme.
Save standard Biogeme YAML results
Description
Save standard Biogeme YAML results
Usage
save_results(fit, filename)
Arguments
fit |
A |
filename |
Destination YAML file. |
Value
fit, invisibly.
Create a segmented parameter expression
Description
The generated native parameters follow Biogeme's public naming contract:
<base>_ref for the reference and <base>_diff_<segment> for differences.
Usage
segment_beta(beta, segmentations, prefix = "segmented")
segmented_beta(beta, segmentations, prefix = "segmented")
Arguments
beta |
A |
segmentations |
A non-empty list of |
prefix |
Native segmentation expression prefix. |
Value
A Biogeme expression compiled through native Segmentation.
Simulate named native Biogeme expressions at fixed estimates
Description
Simulate named native Biogeme expressions at fixed estimates
Usage
simulate(model, expressions = NULL, beta, control = NULL, database = NULL)
Arguments
model |
A |
expressions |
Optional named list of expressions. When omitted, the
model's |
beta |
A |
control |
Optional |
database |
Optional scenario database. It may be a
|
Details
expressions is a named list of complete symbolic expressions. When it is
omitted, the model's simulations list is used. Native Biogeme evaluates
these expressions at the supplied parameter values and returns ordinary R
data frames.
Value
A biogeme_simulation object containing a data frame of values.
Examples
## Not run:
simulation_model <- biogeme_model(
database,
formula = choice_log_probability,
simulations = list(probability = choice_probability)
)
simulated <- simulate(
simulation_model,
beta = fit,
control = biogeme_control(output_directory = tempfile("rbiogeme-sim-"))
)
as.data.frame(simulated)
## End(Not run)
Simulate formulas over Bayesian posterior draws
Description
Evaluates named simulation formulas over native Bayesian posterior draws and returns the native mean and quantile summaries without exposing PyMC objects.
Usage
simulate_bayesian(model, bayesian_results, expressions = NULL,
percentage_of_draws_to_use = 10, lower_quantile = 0.025,
upper_quantile = 0.975, control = NULL)
Arguments
model |
A |
bayesian_results |
A |
expressions |
Optional named list of simulation expressions. |
percentage_of_draws_to_use |
Percentage of posterior draws evaluated by native Biogeme. |
lower_quantile |
Lower posterior-simulation quantile. |
upper_quantile |
Upper posterior-simulation quantile. |
control |
Optional |
Value
A biogeme_simulation object containing native summaries.
Evaluate one native Biogeme formula as an aggregated scalar
Description
This operation delegates to Biogeme's public
calculate_single_formula_from_expression evaluator. Unlike simulate(),
it returns one native aggregated value, which is useful for panel
likelihood formulas.
Usage
simulate_single_formula(
model,
expression,
beta,
number_of_draws,
seed = NULL,
numerically_safe = FALSE,
use_jit = TRUE
)
Arguments
model |
A |
expression |
A Biogeme expression to evaluate. |
beta |
A |
number_of_draws |
Positive number of native Monte Carlo draws. |
seed |
Optional temporary native draw seed. |
numerically_safe |
Whether to request native numerically safe formulas. |
use_jit |
Whether to use native JAX just-in-time compilation. |
Value
One numeric scalar returned by native Biogeme.
Square root of a Biogeme expression
Description
Square root of a Biogeme expression
Usage
## S3 method for class 'biogeme_expression'
sqrt(x)
Arguments
x |
A scalar or Biogeme expression. |
Value
A Biogeme expression.
Build the segmented linear-utility Swissmetro specification from b01b
Description
This uses native LinearUtility and Segmentation during bridge
compilation. The generated parameter names are asc_*_ref and
asc_*_diff_*, matching the Python example.
Usage
swissmetro_b01b_model(
database,
user_notes =
paste0("Example of a logit model with three alternatives: Train, Car and ",
"Swissmetro. Same as 01logit and introducing LinearUtility and ",
"automatic segmentation of parameters.")
)
Arguments
database |
A database prepared by |
user_notes |
Optional notes stored in the native estimation result. |
Value
A biogeme_logit_model.
Prepare the Swissmetro database used by Biogeme's examples
Description
The transformation is deliberately expressed as native database
operations. It removes rows with PURPOSE other than 1 or 3 and rows with
CHOICE == 0, then defines the cost, availability, and /100 scaled
variables used by plot_b01a_logit.py. Row order and original row IDs are
retained. Set panel = TRUE to declare ID as the panel identifier.
Usage
swissmetro_data(
data,
name = "swissmetro",
panel = FALSE,
filter_purpose = TRUE
)
prepare_swissmetro(
data,
name = "swissmetro",
panel = FALSE,
filter_purpose = TRUE
)
Arguments
data |
Numeric Swissmetro data frame. |
name |
Database name. |
panel |
Whether to declare the |
filter_purpose |
Whether to retain only observations with |
Value
A biogeme_database specification.
Build the baseline Swissmetro MNL specification
Description
Build the baseline Swissmetro MNL specification
Usage
swissmetro_mnl_model(database)
Arguments
database |
A database prepared by |
Value
A biogeme_logit_model.
Validate a model using native Biogeme cross-validation
Description
Validate a model using native Biogeme cross-validation
Usage
validate(model, fit, folds = 5L, groups = NULL, seed = NULL, control = NULL)
Arguments
model |
A |
fit |
A |
folds |
Number of validation folds. |
groups |
Optional grouping column. |
seed |
Optional seed controlling native random fold assignment. |
control |
Optional |
Value
A native-bridge validation result.
Validate a data frame for use as a Biogeme database
Description
Validate a data frame for use as a Biogeme database
Usage
validate_biogeme_data(data)
Arguments
data |
An R data frame. |
Value
A validated, copied data frame.
Validate a model before estimation
Description
Model construction, database operations, and native Biogeme model creation are checked without running an optimizer or writing estimation results. The returned diagnostics are ordinary R values; the native model object is created and discarded inside the bridge.
Usage
validate_model(model, control = NULL)
Arguments
model |
A |
control |
Optional |
Details
This is specification validation, not cross-validation. Use validate()
after estimation when the goal is out-of-sample fold evaluation.
Value
An object of class biogeme_model_validation containing the native
validation status, database information, formula names, and parameter
count.
Examples
## Not run:
database <- biogeme_database(
"demo",
data.frame(choice = c(1, 2), x = c(1, 2))
)
model <- logit_model(
database,
choice = "choice",
utilities = list(`1` = 0, `2` = biogeme_beta("b") * variable("x"))
)
validate_model(model)
## End(Not run)
Create a Biogeme data variable
Description
Create a Biogeme data variable
Usage
variable(name)
Arguments
name |
Column name in the Biogeme database. |
Details
variable() creates a symbolic reference. It does not extract an R column
or calculate a vector immediately. Use it in arithmetic and model
expressions after the corresponding database column has been defined.
Value
A variable expression.
Examples
x <- variable("income")
b <- biogeme_beta("b_income")
b * x + 1