Skip to content

Latest commit

 

History

29 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

dMod2 - Dynamic Modeling and Parameter Estimation in R

R-CMD-check

dMod2 builds ODE models of reaction networks, generates and compiles C++ code for them, and estimates their parameters from data. The framework follows the paradigm that derivative information should be used whenever possible: every object it produces carries its symbolic derivatives, so prediction, observation, parameter transformation and objective functions all deliver gradient and Hessian along with their value.

Models are assembled from four kinds of function objects that compose with * (concatenation, chain rule) and + (direct sum over experimental conditions):

Object Constructor Role
odemodel, prdfn odemodel(), Xs() the compiled ODE and its sensitivities
obsfn Y() observation and error functions
parfn P() parameter transformations
objfn normL2(), constraintL2() objective functions

System requirements

Model code is generated and compiled at runtime, so C and C++ compilers are required. On Linux they are usually present, Windows users need Rtools. The default backend is cppDE, which writes C++ with first- and second-order sensitivities; cOde with deSolve is available as an alternative through odemodel(backend = "deSolve").

Installation

remotes::install_github("dModverse/dMod2")

The Remotes: field pulls in cppDE and cOde. master carries the core: equations, compiled prediction, objectives, optimisers and profile likelihood. Four layers live on their own branches, each with its own vignette: devel-symmetry (structural identifiability by symmetry detection and model reduction), devel-petab (PEtab and SBML import and export), devel-EM (nonlinear mixed-effects estimation) and devel-bayes (MCMC and SMC).

Example: STAT5 dimerisation after Epo stimulation

The model of Boehm et al. (2014) describes what happens in BaF3-EpoR cells after stimulation with erythropoietin: STAT5A and STAT5B are phosphorylated, form the three dimers ApA, ApB and BpB, are imported into the nucleus and return to the cytoplasm. The data are relative quantities from mass spectrometry, and they ship with the package.

library(dMod2)
library(ggplot2)
set.seed(5555)

# generated sources and the shared object go here, not into the working directory
outdir <- tempfile("boehm")
dir.create(outdir)

Reactions

Reactions are added one at a time and carry their compartment. The Epo stimulus decays exponentially and enters the phosphorylation rates in closed form, so time appears directly in a rate expression. Nuclear species are assigned before they are produced, so that the reactions importing them do not decide which compartment they belong to.

epo <- "1.25e-7*exp(-Epo_degradation_BaF3*time)"

reactions <- eqnlist() |>
  assignCompartment(nucpApA = "nuc", nucpApB = "nuc", nucpBpB = "nuc", volume = "0.45") |>
  addReaction("2*STAT5A", "pApA", paste0("k_phos*", epo, "*STAT5A^2"),
              "STAT5A-STAT5A phosphorylation", compartment = "cyt") |>
  addReaction("STAT5A + STAT5B", "pApB", paste0("k_phos*", epo, "*STAT5A*STAT5B"),
              "STAT5A-STAT5B phosphorylation", compartment = "cyt") |>
  addReaction("2*STAT5B", "pBpB", paste0("k_phos*", epo, "*STAT5B^2"),
              "STAT5B-STAT5B phosphorylation", compartment = "cyt") |>
  addReaction("pApA", "nucpApA", "k_imp_homo*pApA", "pApA import", compartment = "cyt") |>
  addReaction("pApB", "nucpApB", "k_imp_hetero*pApB", "pApB import", compartment = "cyt") |>
  addReaction("pBpB", "nucpBpB", "k_imp_homo*pBpB", "pBpB import", compartment = "cyt") |>
  addReaction("nucpApA", "2*STAT5A", "k_exp_homo*nucpApA", "pApA export", compartment = "nuc") |>
  addReaction("nucpApB", "STAT5A + STAT5B", "k_exp_hetero*nucpApB", "pApB export", compartment = "nuc") |>
  addReaction("nucpBpB", "2*STAT5B", "k_exp_homo*nucpBpB", "pBpB export", compartment = "nuc") |>
  setCompartmentVolume(cyt = "1.4")

as.eqnvec() turns the table into the ODE right-hand side. The compartment volumes appear as the ratios that convert a flux from the frame it was written in into the frame of the species it acts on.

as.eqnvec(reactions)
## Idx   Inner <- Outer
##   1 nucpApA <- 1*(k_imp_homo*pApA)*(1.4/0.45)-1*(k_exp_homo*nucpApA)
##   2 nucpApB <- 1*(k_imp_hetero*pApB)*(1.4/0.45)-1*(k_exp_hetero*nucpApB)
##   3 nucpBpB <- 1*(k_imp_homo*pBpB)*(1.4/0.45)-1*(k_exp_homo*nucpBpB)
##   4    pApA <- 1*(k_phos*1.25e-7*exp(-Epo_degradation_BaF3*time)*STAT5A^2)-1*(k_imp_homo*pApA)
##   5    pApB <- 1*(k_phos*1.25e-7*exp(-Epo_degradation_BaF3*time)*STAT5A*STAT5B)-1*(k_imp_hetero*pApB)
##   6    pBpB <- 1*(k_phos*1.25e-7*exp(-Epo_degradation_BaF3*time)*STAT5B^2)-1*(k_imp_homo*pBpB)
##   7  STAT5A <- -2*(k_phos*1.25e-7*exp(-Epo_degradation_BaF3*time)*STAT5A^2)-1*(k_phos*1.25e-7*exp(-Epo_degradation_BaF3*time)*STAT5A*STAT5B)+2*(k_exp_homo*nucpApA)*(0.45/1.4)+1*(k_exp_hetero*nucpApB)*(0.45/1.4)
##   8  STAT5B <- -1*(k_phos*1.25e-7*exp(-Epo_degradation_BaF3*time)*STAT5A*STAT5B)-2*(k_phos*1.25e-7*exp(-Epo_degradation_BaF3*time)*STAT5B^2)+1*(k_exp_hetero*nucpApB)*(0.45/1.4)+2*(k_exp_homo*nucpBpB)*(0.45/1.4)

Prediction, observation and error model

Only relative quantities were measured, mixed by the isotope ratio specC17. The error model contributes one standard deviation per observable, estimated alongside the dynamic parameters.

model <- odemodel(reactions, modelname = "boehm_ode", compile = FALSE, outdir = outdir)
x <- Xs(model, 
        optionsOde = list(atol = 1e-8, rtol = 1e-6, maxattemps = 100L, maxsteps = 1e6), 
        optionsSens = list(atol = 1e-8, rtol = 1e-6, maxattemps = 100L, maxsteps = 1e6))

observables <- eqnvec(
  pSTAT5A_rel = "(100*pApB + 200*pApA*specC17)/(pApB + STAT5A*specC17 + 2*pApA*specC17)",
  pSTAT5B_rel = "-(100*pApB - 200*pBpB*(specC17 - 1))/((STAT5B*(specC17 - 1) - pApB) + 2*pBpB*(specC17 - 1))",
  rSTAT5A_rel = "(100*pApB + 100*STAT5A*specC17 + 200*pApA*specC17)/(2*pApB + STAT5A*specC17 + 2*pApA*specC17 - STAT5B*(specC17 - 1) - 2*pBpB*(specC17 - 1))"
)

errors <- eqnvec(
  pSTAT5A_rel = "sd_pSTAT5A_rel",
  pSTAT5B_rel = "sd_pSTAT5B_rel",
  rSTAT5A_rel = "sd_rSTAT5A_rel"
)

g <- Y(observables, x, modelname = "boehm_obs", attach.input = FALSE,
       compile = FALSE, outdir = outdir)
e <- Y(errors, g, modelname = "boehm_err", attach.input = FALSE,
       compile = FALSE, outdir = outdir)

Parameter transformation

The transformation fixes the initial values, splits the total STAT5 pool by the measured ratio, and puts every estimated parameter on a log10 scale. What comes out is the set of nine parameters the fit sees.

innerpars <- getParameters(model, g, e)
estimated <- c("Epo_degradation_BaF3", "k_exp_hetero", "k_exp_homo", "k_imp_hetero",
               "k_imp_homo", "k_phos", "sd_pSTAT5A_rel", "sd_pSTAT5B_rel", "sd_rSTAT5A_rel")

p <- eqnvec() |>
  define("x~x", x = innerpars) |>
  define("x~0", x = c("pApA", "pApB", "pBpB", "nucpApA", "nucpApB", "nucpBpB")) |>
  insert("STAT5A ~ 207.6*ratio") |>
  insert("STAT5B ~ 207.6 - 207.6*ratio") |>
  insert("ratio ~ 0.693") |>
  insert("specC17 ~ 0.107") |>
  insert("x~exp10(x)", x = estimated) |>
  P(modelname = "boehm_trafo", condition = "Boehm2014",
    compile = FALSE, outdir = outdir)

getParameters(p)
## [1] "k_phos"               "Epo_degradation_BaF3" "k_exp_homo"          
## [4] "k_exp_hetero"         "k_imp_homo"           "k_imp_hetero"        
## [7] "sd_pSTAT5A_rel"       "sd_pSTAT5B_rel"       "sd_rSTAT5A_rel"

Compilation

Every object was built with compile = FALSE, so one call compiles them together into a single shared object. Compile flags are per file: the ODE and its sensitivities are built with OpenMP, the observation and transformation sources without, and the report says so.

compile(g, x, p, e, output = "boehm", cores = 4)
## using C++ compiler: g++
##   every source : -O2 -DNDEBUG -w -fPIC -DKLU -I/home/simon/.cache/R/cppDE/deps-sundials-7.4.0-suitesparse-7.10.0/include/suitesparse
##   2 of 5 also  : -DKLUBTF=0 -DKLUAMD=0 -fopenmp

Data

data(boehm)
head(boehm)
##   time        name     value sigma condition
## 1  0.0 pSTAT5A_rel  7.901073    NA Boehm2014
## 2  2.5 pSTAT5A_rel 66.363494    NA Boehm2014
## 3  5.0 pSTAT5A_rel 81.171324    NA Boehm2014
## 4 10.0 pSTAT5A_rel 94.730308    NA Boehm2014
## 5 15.0 pSTAT5A_rel 95.116483    NA Boehm2014
## 6 20.0 pSTAT5A_rel 91.441717    NA Boehm2014

Objective and multi-start fit

normL2() combines data, the composed prediction g*x*p and the error model into an objective returning value, gradient and Hessian. Two directions of this model are practically non-identifiable, so a weak prior on the dynamic parameters keeps them at finite values. It deliberately stays off the sd_*: reml() below solves the stationarity condition of the data term for those and would not see a prior on them.

pouter <- structure(rep(-1, length(getParameters(p))), names = getParameters(p))
dyn <- setdiff(names(pouter), grep("^sd_", names(pouter), value = TRUE))
data <- as.datalist(boehm)
obj <- normL2(data, g*x*p, e) + constraintL2(pouter[dyn], sigma = 4)

A single trust-region run finds a local optimum. Multi-start search shows how the objective is shaped: fifty fits from randomly drawn starting points, sorted by their value, give the usual staircase with a plateau of runs that reached the best optimum.

fits <- mstrust(obj, center = pouter, sd = 3, fits = 100, cores = 8,
                rinit = 0.1, rmax = 10, iterlim = 5000)
outframe <- as.parframe(fits)

plotValues(outframe, tol = 0.1, value < 1e4)

best <- as.parvec(outframe)

Error model by restricted maximum likelihood

Estimating the standard deviations by maximum likelihood leaves them too small: the residuals have already spent parameters on the mean. reml() corrects for that by charging every data point its own leverage, which is what the restricted likelihood asks for, rather than giving each the same share of the parameter budget.

The restricted likelihood is the likelihood of error contrasts, combinations of the data that carry no information about the mean parameters and therefore leave sigma unbiased (Patterson and Thompson (1971)).

remlfit <- reml(obj, best)

rbind(ML = best[remlfit$errpars], REML = remlfit$argument[remlfit$errpars])
##      sd_pSTAT5A_rel sd_pSTAT5B_rel sd_rSTAT5A_rel
## ML        0.5876483      0.8200724      0.4987661
## REML      0.6254742      0.8442083      0.5248422
remlfit$dof     # n_g minus the leverage each observable spends
## pSTAT5A_rel pSTAT5B_rel rSTAT5A_rel 
##    13.67813    14.22147    14.10040
remlfit$rank    # effective number of mean parameters
## [1] 6

remlLeverage() shows where the budget goes. The rank is what the model actually spends on the mean, not the nominal parameter count.

lev <- remlLeverage(obj, best)
tapply(lev$leverage, lev$name, sum)
## pSTAT5A_rel pSTAT5B_rel rSTAT5A_rel 
##    2.372391    1.742590    1.885019

Fit and uncertainty band

The band is one standard deviation and therefore belongs to the REML estimate. The dynamic parameters move a little as well: the three correction factors differ, which reweights the observables against each other.

pars <- remlfit$argument
times <- seq(0, 240, length.out = 200)

prdout <- (g*x*p)(times, pars, deriv = FALSE) |> as.data.frame(errfn = e)

plot((g*x*p)(times, pars, deriv = FALSE), data) +
  geom_ribbon(data = prdout, aes(x = time, ymin = value - sigma, ymax = value + sigma),
              linetype = "dashed", alpha = 0.3)

Profile likelihood

Profiles re-optimise every other parameter while one is held away from its optimum. They are the identifiability statement the Hessian at the optimum cannot give (Raue et al. (2009)).

Profiling a nuisance parameter out leaves the profile biased as a likelihood. The general correction is Barndorff-Nielsen’s adjusted profile likelihood (1983, 1994), with Cox and Reid (1987) as the tractable approximation. For Gaussian errors with the error model profiled out it reduces to the restricted likelihood above, which is what reml() computes. dMod2 implements that case only, and the effective rank it reports is the p the threshold below is calibrated on.

Three details decide whether the result is read off or guessed.

The profile must reach the threshold the interval is taken at, so profileThreshold() computes it once and profile(), plotProfile() and confint() all use the same number. method = "F" is the finite-sample calibration, exact for Gaussian errors with a single sigma profiled out and the generalisation of the classical t interval.

Everything refers to the objective that was actually optimised, prior included. stop = "value" and val.column = "value" select the total rather than the data term, and only for the total is the fit the minimum, so only there does the line marking the optimum sit where it belongs.

And the profiles are anchored at that fit, because a profile is measured against the value at its origin.

nd     <- sum(vapply(data, nrow, 0L))
thr_95 <- profileThreshold(0.95, method = "F", n = nd, p = remlfit$rank)
thr_90 <- profileThreshold(0.90, method = "F", n = nd, p = remlfit$rank)

profiles <- profile(
  obj, best, names(best), limits = c(-5, 5), cores = length(best), delta = thr_95,
  stepControl = list(stepsize = 1e-4, min = 1e-4, max = Inf,
                     atol = 1e-3, rtol = 1e-3, limit = 200, stop = "data"),
  algoControl = list(reoptimize = TRUE),
  optControl  = list(rinit = 0.1, rmax = 10, iterlim = 20))

plotProfile(profiles, mode %in% c("data", "prior"), threshold = c("90%" = thr_90, "95%" = thr_95))

confint(profiles, level = 0.95, val.column = "data",
        method = "F", n = nd, p = remlfit$rank)
##                                      name      value      lower      upper
## Epo_degradation_BaF3 Epo_degradation_BaF3 -1.5680325 -1.7398858 -1.3903480
## k_exp_hetero                 k_exp_hetero -4.4562477       -Inf -2.9546299
## k_exp_homo                     k_exp_homo -2.2032225 -2.5569151 -1.8895645
## k_imp_hetero                 k_imp_hetero -1.7872354 -1.9108026 -1.6583018
## k_imp_homo                     k_imp_homo  1.4788926  0.1002196        Inf
## k_phos                             k_phos  4.1976370  4.1064208  4.2964734
## sd_pSTAT5A_rel             sd_pSTAT5A_rel  0.5876483  0.4098707  0.7938722
## sd_pSTAT5B_rel             sd_pSTAT5B_rel  0.8200724  0.6648408  1.0126894
## sd_rSTAT5A_rel             sd_rSTAT5A_rel  0.4987661  0.3488114  0.6910599

The plot separates the two terms, and that is where the two non-identifiable directions become visible: for k_imp_homo and k_exp_hetero the data curve stays level and only prior rises, slowly. Nuclear import of the homodimers is already fast enough to be rate-limited by phosphorylation, and the heterodimer barely returns to the cytoplasm within the measured 240 minutes. Neither reaches the threshold inside the searched range, so those intervals stay open on one side. The prior keeps the fit at finite values without pretending to close them.

Citation

There is no publication for dMod2 yet. The paper below describes dMod, its predecessor, and is the reference for the modelling framework:

Kaschek D, Mader W, Fehling-Kaschek M, Rosenblatt M, Timmer J (2019). Dynamic Modeling, Parameter Estimation, and Uncertainty Analysis in R. Journal of Statistical Software, 88(10), 1-32. doi:10.18637/jss.v088.i10

About

Dynamic modelling and parameter estimation in ODE models of reaction networks. Successor to dMod.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages