mohannadibrahim.dev

README.md/case studies/c3 eligibility

Case study 03 / Reliability Production

Keeping an eligibility API fast, observable, and fresh

I owned a C#/.NET eligibility-and-claim-status service at Ameritas: p95 under 200 ms across about 400,000 requests a month, with releases on demand and alarms that catch bad eligibility data in minutes.

At a glancec3 / eligibility

Role
Ameritas intern [CONFIRM: title held while owning this service]
Company
Ameritas
Scale
~400,000 requests / month
Dates
[CONFIRM: dates on this service]
Latency
p95 under 200 ms
Result
days to under 15 min to detect
Stack
C#/.NET, AWS (ECS Fargate, CloudWatch)
On this page
  1. Problem
  2. Constraints
  3. Approach
  4. Decision
  5. Evaluation
  6. Result
  7. What I'd do next

01Problem

The service answers eligibility and claim-status requests, about 400,000 a month. Releases ran on a monthly window, and a class of eligibility mismatch could sit unnoticed for days.

Placeholder[CONFIRM: who calls the service and what a mismatch looks like to them].

02Constraints

  • Latency had to hold: p95 under 200 ms at about 400,000 monthly requests.
  • Eligibility answers must not go stale, so any caching has to respect member-level changes.
  • [CONFIRM: compliance, uptime, or review constraints that applied to this service].

03Approach

Four changes, each aimed at one failure mode.

  1. Hosting. Moved the service to ECS Fargate, which took releases from a monthly window to on-demand.
  2. Detection. CloudWatch metrics and alarms cut the time to detect a class of eligibility mismatch from days to under 15 minutes.
  3. Safety net. Unit-test coverage went from about 40% to 85%.
  4. Caching. A short-TTL cache keyed on member, after a senior engineer flagged stale-eligibility risk in the first design.
ECS FARGATE TRAFFIC Requests eligibility, claim status ~400,000 / month SERVICE Eligibility + claim status C#/.NET, p95 under 200 ms CACHE Keyed on member short TTL after a stale-eligibility flag metrics + alarms OBSERVABILITY CloudWatch mismatch caught in under 15 min TRAFFIC Requests, ~400,000 / month ECS FARGATE SERVICE Eligibility + claim status C#/.NET, p95 under 200 ms CACHE Keyed on member short TTL after a stale-eligibility flag metrics + alarms OBSERVABILITY CloudWatch mismatch caught in under 15 min
Fig. 1The service runs on ECS Fargate with a member-keyed short-TTL cache in front of its lookups. CloudWatch metrics and alarms watch it, which is what brought detection time from days to under 15 minutes. [CONFIRM: exact cache placement relative to the data store].

04Decision

A short-TTL cache keyed on member. I revised the caching design after a senior engineer flagged stale-eligibility risk. Keying on the member and keeping the TTL short bounds how long a changed eligibility answer can be wrong.

Alternatives considered
The earlier caching designrevised

A senior engineer flagged stale-eligibility risk. [CONFIRM: what the original cache key and TTL were].

Monthly release windowreplaced

Moving to ECS Fargate allowed on-demand releases instead. [CONFIRM: any other hosting options weighed].

Short-TTL cache keyed on memberchosen

Keeps responses fast without serving another member's data or a long-stale answer.

05Evaluation

Unit-test coverage went from about 40% to 85%. CloudWatch metrics and alarms track the service in production.

Placeholder[CONFIRM: how p95 was measured and over what window, and which tests cover the cache behavior].

06Result

Days to <15 mintime to detect a class of eligibility mismatch

p95 under 200 ms across about 400,000 monthly requests, releases on demand instead of monthly, and unit-test coverage from about 40% to 85%.

07What I'd do next

  • Alarm on cache staleness directly. Today the alarms catch a mismatch class; a staleness metric would catch cache drift before a member does. [CONFIRM: whether this already exists].
  • Cover the cache with tests. Pin the member key and TTL behavior so the original risk cannot return quietly.
  • Publish the latency trend. Keep p95 visible against the 200 ms target on a dashboard.