Skip to content
← Back to GG Insights

Case Study

Independent Validation of a Term Life Model Operating Across Six Valuation Bases for a Fortune 500 Life Insurer

9 min readPublished August 2026Graeme Group

Engagement Type

Fixed Fee per Model – Comprehensive Model Validation

Practice Area

Life Insurance – Statutory, PBR & Multi-Basis Model Risk

Jurisdiction

USA

Objective

Graeme Group was retained by a Fortune 500 US life insurer to perform a comprehensive independent validation of the ALFA actuarial projection model used to value its term life block, covering inputs, calculations, outputs and governance.

What made this model unusual was the breadth of what a single model is asked to do. It supports six distinct valuation bases simultaneously, legacy statutory reserving, principle-based reserving under VM-20, US GAAP under the long-duration targeted improvements standard, an offshore economic basis, cash flow testing, and multi-year forecasting. Each basis has its own prescribed assumptions, its own regulatory logic, and its own definition of a correct answer. One model has to satisfy all six without accommodations made for one basis distorting another.

Compounding this, the term model is a fork of a shared library that also drives several other product lines, so a variable customized for one product can propagate into others. Roughly a third of the variable definitions had been customized away from vendor defaults, and a material subset of those touched term business directly or through shared code.

The model was classified at the Company’s highest internal risk rating, driven by the size of the reserves it produces, the number of external reporting processes consuming its output, and its structural complexity. It had not been comprehensively validated for several years, the previous exercise having been a conversion validation performed when the block was migrated onto the platform. The Company wanted evidence, rather than assertion, that the model was fit for use on every basis it supports.

Scope of Work

The validation covered the full term life block, base term and guaranteed issue term, across all rider structures, underwriting classes and smoking statuses, spanning issue years reaching back nearly three decades and five distinct product generations. More than 400,000 in-force contract records were in scope, valued as at the year-end anchor date.

Coverage was end-to-end on all six bases: input and assumption integrity, the soundness of the calculations, the completeness and downstream referencing of outputs, and the governance framework surrounding the model. Procedures were applied consistently across bases rather than concentrated on the largest, on the principle that the least-scrutinized basis is usually where the risk has accumulated.

The engagement opened with a written validation plan defining 25 discrete procedures across five categories, each with a stated purpose, scope and nominated workpaper, agreed with the model owner before fieldwork began, so the basis of every subsequent conclusion was fixed in advance rather than negotiated after the fact.

One boundary was set deliberately: the underlying asset model feeding the cash flow testing basis was excluded. Scenario integrity, the liability-side projection logic and the integration points where asset output enters the term model were all validated, but the asset model itself was out of scope and the report says so plainly.

Services Provided

Input and Assumption Integrity

Every assumption family feeding the model was extracted from source and reconciled cell by cell against the values held in the model: mortality and mortality improvement, lapse and post-level-term persistency, expenses, term conversion rates, premium rates and interest rates. This detailed reconciliation work is where most of the findings originated. More than 170 production premium rate tables were reconciled against their source workbooks, along with the full population of lapse tables drawn from sixteen separate source files.

Separately, the in-force extract loaded into the model was reconciled against the upstream data warehouse extract on both count and total controls, and a 27-test data integrity framework was applied to field-level coherence, policy uniqueness, blank and default values, values outside documented product boundaries, and consistency between related fields.

Independent Recalculation

The core of any calculation validation is whether the model’s answer can be reproduced without using the model. We built independent recalculators for every basis in scope, a suite of Excel/VBA calculators constructed from first principles, and a separate independent projection engine written in Python.

The governing rule on both was strict: the independent side is assembled from its own formulas, never by copying the model’s. The model’s resolved assumption inputs may be fed in, but the assembly, the projection mechanics and the reserve formulas are the challenger’s own. Where an independently derived cash flow legitimately differs, the difference is documented and explained rather than tuned away. Every parameter in the Excel calculators is supported by a citation to an independent authority, a regulation, an actuarial standard, or a client assumption memorandum that is itself independently sourced. Model files are not accepted as support for the output they are intended to test.

Sample selection used a gross reserve-weighted stratified random approach: policies were partitioned into cohorts defined by the concatenation of the key reserve drivers, and cohorts selected to guarantee coverage of material exposures and the full range of contractual variation, rather than a simple random draw that would over-represent small, routine policies.

Conceptual Soundness and Code Review

Beyond arithmetic, we assessed whether each basis was methodologically right: whether the reserve mechanics suited the contractual features present in the book, whether the interaction between bases was coherent, and whether cohort construction under the GAAP basis grouped business in a way the standard supports. Customized variables were sample-tested against the Company’s own coding guidelines, and seventeen change management packages released ahead of the anchor date were reviewed alongside an enterprise regression between the two most recent model versions.

Output, Downstream and Governance Review

Model output was traced through to the reporting artifacts that consume it, and a year-over-year reasonableness assessment attributed movements to identifiable drivers rather than accepting them as noise. On governance, the model was tested against the Company’s own published standards, ownership, documentation, inventory accuracy, assumption governance, change control, environment segregation, storage and security, roles and independence, and validation cadence. Testing each requirement against the client’s own written standard, rather than a generic checklist, is what makes governance findings actionable instead of arguable.

Findings Architecture: Why We Do Not Assign Severity Ratings

Findings on this engagement are not assigned severity ratings. They are classified on two axes only: the domain where the root cause sits (inputs, calculations, outputs or governance), and the type of finding, an Issue, where numbers demonstrably do not tie; a Simplification, where the model intentionally reduces detail relative to source or specification; a Conceptual Soundness point, where the methodology is questionable on its merits irrespective of implementation; or a Maintenance and Documentation point, where the derivation, evidence or write-up is missing.

This is a considered design choice, tested against the conventional alternative. A high/medium/low label attached to an individual observation pre-judges materiality before the report-level materiality framework has been applied to it. In practice it also shifts the discussion: reviewers negotiate the label rather than the substance. And it goes stale, because a finding’s significance changes as remediation proceeds while the label does not. A rounding mismatch is still an Issue, because the numbers either tie or they do not. The observation log records what was observed, against source, with a hyperlink to the evidence. It is an evidentiary record rather than a dashboard, and it is kept monochrome, with no traffic-light ratings, for that reason. Materiality is assessed once, at report level, where the materiality framework applies.

Deliverables

The engagement resulted in:

  • A comprehensive Model Validation Report concluding that the model was validated for use across all six bases supported, subject to the findings raised, with a documented conclusion for each of the four validation domains.
  • A written Validation Plan setting out all 25 procedures, agreed before fieldwork, and reusable as the specification for the next validation cycle.
  • An observation log of 23 findings, each classified by domain and type, traced to its supporting evidence, and written to drop straight into the Company’s remediation process.
  • A structured library of 39 numbered workpapers across the input, calculation, output and governance series, with a legend mapping each to the procedures and findings it supports.
  • Independent recalculators for every basis in scope, Excel/VBA calculators plus a modular Python projection engine, transferred as a permanent challenger capability re-runnable in later cycles without external assistance.
  • Reusable reference artifacts retained by the Company: a cross-basis assumption inventory, a product feature grid spanning all five generations, and an end-to-end model flow diagram.
  • Sign-off by qualified actuarial reviewers.

Outcome

  • Multi-Basis Assurance: Evidenced comfort that a single model serving six regulatory and reporting regimes performs appropriately on each, not an opinion on the largest basis with the others assumed to follow.
  • Reconciliation on the Foundations: The in-force population loaded into the model reconciled to the upstream extract with no difference on either count or total controls, so every downstream reserve conclusion rested on a complete population.
  • A Genuine Independent Challenger: Recalculators built from first principles reproduced the model’s results within tolerance across the sampled batches, one basis tying to zero difference on every cash flow compared. Built from independent formulas and independently cited parameters, their agreement is genuine evidence rather than a restatement of the model’s own logic.
  • Findings That Can Be Acted On: Twenty-three findings across all four domains, weighted toward inputs and assumptions. Each identifies a root cause rather than a symptom, and its classification tells the model owner what kind of work closes it, a correction, a documented rationale, a methodology decision, or a piece of missing evidence.
  • Governance Tested Against the Company’s Own Standards: Testing against the Company’s published model risk standards rather than a generic framework produced governance findings that map directly onto obligations the Company had already set for itself.
  • Scope Limits Reported Plainly: Where an independent recalculation route could not be fully established within the engagement window, that was reported as such, with the root cause identified, rather than presented as a clean result.

By building genuine independent challengers for every basis, testing assumptions back to source at cell level, and reporting findings in a structure built for remediation rather than for a dashboard, our team left behind a validation plan, a workpaper library and a challenger toolkit that materially reduce the cost of the next validation cycle.


Next step

Facing something similar? Discuss a project.

Contact Graeme Group

Next step

Discuss a project.

Tell us what you are validating, building, or staffing.