Select a meeting-summary pilot — decision exercise
Exercise: select a meeting-summary pilot
Original fictional scenario. No cloud account, API calls or paid services are required. Difficulty: beginner · Estimated duration: 10–15 minutes
A team needs audio plus text input and text summaries. It requires an approved deployment, at least 90 accepted outputs per 100 reviewed meetings, and a 95th-percentile response time of at most 4 seconds. All quoted costs below assume the same forecast volume and are invented.
| Candidate | Audio input | Deployment approved | Accepted / 100 | Time | Monthly estimate |
|---|---|---|---|---|---|
| Juniper | Yes | Yes | 94 | 3 sec | USD 80 |
| Hawthorn | No | Yes | 96 | 2 sec | USD 50 |
| Rowan | Yes | Yes | 97 | 6 sec | USD 65 |
| Larch | Yes | Unresolved | 95 | 2 sec | USD 60 |
Your decision
- Identify eligible candidates and name each failed or unresolved condition.
- Recommend the next pilot candidate using the supplied measurements.
- Describe the additional evidence you would collect before production.
- Recalculate if the hard time limit changes to 7 seconds. State what remains fixed.
Rubric (10 house points)
- 4 points: correctly classify each candidate with its specific reason.
- 2 points: choose Juniper under the original requirements.
- 2 points: request representative workload/reviewer evidence and deployment checks.
- 2 points: reassess Rowan under the new time limit while retaining the other gates.
Reference solution
Juniper meets the original stated requirements. Hawthorn lacks required audio input; Rowan misses the four-second limit; Larch has an unresolved mandatory deployment decision. Advance Juniper to the next controlled pilot stage. This table does not establish production reliability.
If the limit changes to seven seconds, Rowan becomes eligible under the other unchanged premises. Its higher score and lower forecast cost make it the preferred candidate in this exercise. The change does not create audio capability for Hawthorn or approve Larch’s deployment.
Before production, validate the test set’s relevance, scoring consistency, peak-load behavior, chosen service configuration, current rates and measured token usage. Record the decision owner, model version and evidence date.
Sources and scope
All candidate names, numbers, thresholds, rubric and conclusions above are house-authored. General selection and evaluation principles are grounded in:
- https://docs.cloud.google.com/docs/ai-ml/generative-ai/develop-generative-ai-application
- https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/evaluation-overview
- https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricing
- https://docs.cloud.google.com/architecture/framework/perspectives/ai-ml/reliability