README.md/case studies/c3 eligibility
Case study 03 / Reliability Production
Keeping an eligibility API fast, observable, and fresh
I owned a C#/.NET eligibility-and-claim-status service at Ameritas: p95 under 200 ms across about 400,000 requests a month, with releases on demand and alarms that catch bad eligibility data in minutes.
At a glancec3 / eligibility
- Role
- Ameritas intern [CONFIRM: title held while owning this service]
- Company
- Ameritas
- Scale
- ~400,000 requests / month
- Dates
- [CONFIRM: dates on this service]
- Latency
- p95 under 200 ms
- Stack
- C#/.NET, AWS (ECS Fargate, CloudWatch)
- Tested with
- Unit tests, ~40% to 85% coverage
On this page
01Problem
The service answers eligibility and claim-status requests, about 400,000 a month. Releases ran on a monthly window, and a class of eligibility mismatch could sit unnoticed for days.
Placeholder[CONFIRM: who calls the service and what a mismatch looks like to them].
02Constraints
- Latency had to hold: p95 under 200 ms at about 400,000 monthly requests.
- Eligibility answers must not go stale, so any caching has to respect member-level changes.
- [CONFIRM: compliance, uptime, or review constraints that applied to this service].
03Approach
Four changes, each aimed at one failure mode.
- Hosting. Moved the service to ECS Fargate, which took releases from a monthly window to on-demand.
- Detection. CloudWatch metrics and alarms cut the time to detect a class of eligibility mismatch from days to under 15 minutes.
- Safety net. Unit-test coverage went from about 40% to 85%.
- Caching. A short-TTL cache keyed on member, after a senior engineer flagged stale-eligibility risk in the first design.
04Decision
A short-TTL cache keyed on member. I revised the caching design after a senior engineer flagged stale-eligibility risk. Keying on the member and keeping the TTL short bounds how long a changed eligibility answer can be wrong.
A senior engineer flagged stale-eligibility risk. [CONFIRM: what the original cache key and TTL were].
Moving to ECS Fargate allowed on-demand releases instead. [CONFIRM: any other hosting options weighed].
Keeps responses fast without serving another member's data or a long-stale answer.
05Evaluation
Unit-test coverage went from about 40% to 85%. CloudWatch metrics and alarms track the service in production.
Placeholder[CONFIRM: how p95 was measured and over what window, and which tests cover the cache behavior].
06Result
p95 under 200 ms across about 400,000 monthly requests, releases on demand instead of monthly, and unit-test coverage from about 40% to 85%.
07What I'd do next
- Alarm on cache staleness directly. Today the alarms catch a mismatch class; a staleness metric would catch cache drift before a member does. [CONFIRM: whether this already exists].
- Cover the cache with tests. Pin the member key and TTL behavior so the original risk cannot return quietly.
- Publish the latency trend. Keep p95 visible against the 200 ms target on a dashboard.