Case Study
Independent Validation of a Fixed Annuity Valuation Model for a Fortune 500 US Life Insurer
Objective
Graeme Group was engaged by a Fortune 500 US life insurer to perform an independent validation of the AXIS model used to value its fixed annuity block. The engagement was commissioned by the Company’s second-line model risk function, with the model owner sitting in annuity valuation, and was structured to produce a report that both functions could sign.
The model is a single valuation engine covering two distinct categories of business. On one side sit payout annuities, immediate, deferred income, life contingent and term certain, where the liability is a stream of guaranteed payments and the reserve turns on mortality, the prescribed valuation rate, and the contractual payment form. On the other sit single premium deferred annuities, where the reserve turns on account value mechanics, guaranteed crediting floors and product-specific surrender charge schedules. The valuation approach differs materially between the two, and both are produced from the same model on the same reporting cycle.
The model was classified at the Company’s highest internal risk rating. That rating reflected the materiality of the statutory reserves it produces and the fact that its output is booked directly to the general ledger, rather than any exotic quantitative content, the reserving methodologies involved are well established and well understood. The Company wanted independent evidence on three questions: whether the data flowing into the model is complete and faithful to source, whether the reserves it produces can be reproduced by a party that did not build it, and whether the governance around it meets the standards the Company has set for a model of that rating.
Scope of Work
The validation covered the in-scope payout and deferred annuity products as at the mid-year anchor date, on both the statutory and tax reserving bases. US GAAP was formally out of scope, as the reporting entity does not report on that basis, a boundary agreed and documented rather than left implicit. A defined set of immaterial products and contract variants was also excluded, with the materiality rationale recorded in the report.
Within that perimeter the engagement was deliberately end-to-end, following the data flow rather than functional boundaries. It began upstream of the model, at the extracts arriving from the Company’s policy administration systems, and followed them through the data import and mapping layer, the model’s cell and table structures, the batch processes that run the valuation, the seriatim output, and finally into the ledger posting file that books the reserve. The governance wrapper was assessed alongside: documentation, model inventory, change management, environment segregation, access control, escalation and ongoing monitoring.
Two components were explicitly outside the perimeter and named as such: downstream tools that consume model output for roll-forward, projection and reinsurance purposes, and the derivation of the actuarial assumptions themselves. The validation tested that assumptions were faithfully implemented and consistently applied, not whether they were the right assumptions to have chosen.
Services Provided
Input Validation
The in-force data loaded into the model was reconciled against the upstream source extracts, and a battery of data quality tests was applied to the loaded population: policy record uniqueness, blank and default field assessment, reconciliation of demographic splits back to total lives, and a randomly selected spot-check sample per block traced field by field to source.
The data import and mapping layer received particular attention, because it is the point where a silent error does the most damage. Rather than reading the mapping logic and forming a view, we independently reproduced the conversion in Excel and reconciled our output against the model’s. The mapping code itself was then reviewed directly within the platform for structure, nesting depth and readability, a transformation that produces the right answer today but cannot be safely modified is still a finding.
Calculation Validation
Reserves were independently recalculated at seriatim level on a representative policy sample. Sampling used a gross reserve-weighted stratified random approach: policies were partitioned into cohorts defined by the concatenation of the key reserve drivers, product type, life basis, gender and reserving approach, and the largest cohorts were selected such that cumulative coverage exceeded 80% of the total block reserve before any random selection took place within them. This concentrates testing where reserve materiality sits while still exercising the full range of contractual forms, and it produces a sample whose coverage can be demonstrated rather than asserted.
Each sampled policy was recalculated in independent testware built outside the model, from the prescribed methodology and independently sourced valuation parameters, and compared against the model’s own seriatim output against a defined tolerance. Differences were investigated to root cause and reported as findings where they were genuine, rather than absorbed into a tolerance band.
Alongside the sample work, we ran a consistency review across the model’s cell structure using platform queries to compare assumption application across the statutory basis, the tax basis and the reserving approach, the test that identifies a policy segment configured differently from its peers. The model was then re-run end to end from the production dataset, batch logs reviewed for errors and warnings, and the aggregated output reconciled to the extracts used for reporting.
Output and Downstream Review
Model output does not stop at the model. We traced the seriatim extract into the ledger posting file that books the reserve, checking sign conventions, product line mappings and the line-by-line allocation to sub-product codes, and reviewed the evidence for the reconciliation control operating over that handoff. This is where a technically correct model can still produce an inaccurate reported balance, and it is routinely outside the scope of a validation that stops at the model boundary.
Governance Assessment
Governance was assessed against the Company’s own model risk framework rather than a generic checklist, across the documentation set, the model inventory entry, change management, deployment and environment practices, access and security, the escalation route for issues found in production use, and the ongoing monitoring controls. Change management was tested by sampling actual changes released during the period and following each through to its approval evidence, confirming that implementation and approval were performed by different individuals. Individual risks were scored on a defined scale, so that the governance conclusion rests on a visible set of component assessments rather than a single overall impression.
Model Validation Report
The engagement concluded with a Model Validation Report delivered on the Company’s own reporting template and structured to serve both the model owner and the second-line function. Findings were separated by type, genuine issues, intentional simplifications, improvement opportunities and inconsequential observations, because those four categories call for four different responses, and collapsing them into a single ranked list obscures which is which. Each was accompanied by a description of the underlying risk, a recommendation, and a proposed remediation owner and date.
Deliverables
The engagement resulted in:
- A Model Validation Report issued on the Company’s own template, setting out an overall model rating, a per-domain assessment across inputs, calculations, outputs and governance, and a full findings register with recommendations, owners and target dates.
- Independent recalculation testware for both the payout and deferred annuity lines, extended during the engagement so that a materially broader set of policies can be recalculated independently, and transferred to the Company for reuse.
- A documented sampling methodology, including the cohort definitions and the coverage analysis evidencing that the sample represented the block.
- A draft assumptions inventory capturing the key actuarial assumptions consumed by the model, an artifact the Company did not previously hold in consolidated form.
- A supporting workpaper set covering the input and calculation testing, handed over with the report.
- Sign-off by qualified actuarial reviewers, with formal acknowledgement by the model owner and approval by the second-line model risk function.
Outcome
- Reserve Confidence: Independent seriatim recalculation confirmed that reserves for the sampled population were reproducible within tolerance from outside the model, giving the Company evidenced comfort in the figures being booked and reported.
- A Demonstrable Sample: Because sampling was reserve-weighted and cohort-based with documented coverage, the Company can show a reviewer exactly what proportion of the block the testing spoke to, a materially stronger position under audit or regulatory challenge than a simple random sample of comparable size.
- An Assumption Long Overdue for Examination: The validation identified a deliberate modeling simplification that had been approved years earlier and had not been revisited since. It remained defensible, but its effect had never been quantified. Sizing it gave the Company a current, evidenced basis for deciding whether to retain it.
- Upstream Data Corrected: Field-level testing of the loaded population surfaced a small set of records holding an incorrect value originating in an upstream system, outside the model itself, the kind of defect that a model-boundary validation never sees.
- Governance Gaps With Named Remedies: The governance assessment identified specific, closable weaknesses, each mapped to the Company’s own standard and each given an owner and a target date so that findings entered the remediation process rather than a report appendix.
- A Right-Sized Risk Assessment: Where testing supported a lower assessment of a component risk than recorded in the inventory, the lower assessment was proposed and evidenced. Validation should be capable of reducing a rating as well as raising one; otherwise the rating stops carrying information.
- A Continuing Mandate: The Company subsequently engaged Graeme Group to validate further models in its inventory, on a materially larger scope.
This engagement shows Graeme Group validating a model whose difficulty lies not in mathematical exoticism but in breadth, materiality and operational reality, two distinct product families, two reporting bases, a data path running from policy administration through to the ledger, and reserves booked straight to the accounts. By reproducing the data conversion independently rather than reviewing it, sampling in a way whose coverage can be evidenced, and following the numbers past the model boundary into the ledger, our team gave both the model owner and the second line a report they could each stand behind.