Simara Architecture Studio

Before you build,
make the decisions clear.

Map the work. Design the handoffs. Put a cost and a test beside the proposal.

AI interpretsSoftware enforcesPeople decide

Define the work

Describe a real process.

Use a concrete example. Include what starts the work, what decisions follow, and what must happen at the end.

Keep the whole method in view

Work through the three sections below. A blank answer remains an open decision. The first proposal gives you a capability map; return here to confirm how its AI and human steps are costed.

Scope the work: four steps

1. Business requirement to separate capabilities

The proposal breaks the process into individual tasks with inputs, outputs and owners.

2. Assign AI, company systems and people

The four technical questions below determine where AI fits and which controls it needs.

3. Define the limits of the proposed system

Record input limits, quality, access, freshness and response time. For response time, use the time within which 95% of requests must finish. An average is not the same measure.

4. Carry the limits into the proposed scope

The report carries the capabilities, exclusions, assumptions, boundaries and acceptance tests together. It remains a proposed scope until agreed.

Size the use case: four steps

1. Establish demand and count the actual work

After the first proposal, map the resulting AI and human capabilities here. Until then, the budget and ROI remain incomplete.

2. Set token budgets for the distribution of requests

The workload table below sets shares, total input, output, calls and retries. Break down each input here, including the largest case.

Typical request

These three parts must add up to 3,000 input tokens in the workload table.

Cache write costs and reuse

Provider behaviour differs. Cached reads are priced separately from initial writes. Enter no cache savings until the route and reuse pattern are understood.

3. Price the full operation and assess alternatives

Model prices come from the selected target model. Add all retained staff work, including approvals and exception handling. Group activities performed together once.

Could requests wait for batch processing?

4. Test the fragile assumptions

The report recalculates double demand and a more demanding input distribution, with cold cache, more repeated calls and more staff effort. It shows when a scenario crosses the budget ceiling.

Business value and ROI: four steps

1. Establish today's measured baseline

An assumed baseline can show a what-if calculation. It leaves the measured-baseline step open and cannot establish savings or investment payback.

2. Compare the future state in the same units

Address all five value pillars

Explain the expected contribution, or state why a pillar is not relevant. Different kinds of value should remain visible even when they cannot be priced.

3. Reuse the full recurring cost from sizing

4. Show payback and sensitivity

The calculator keeps build cost separate. It shows payback only for a positive cash case, and tests both half the demand and half the gain per request. Technical feasibility and investment value remain separate decisions.

Technical feasibility, in four questions

These are the course assessment questions. Answer each for this use case, naming differences between capabilities. Enter “Unknown” where an answer is missing. Studio must retain that gap and show how each answer affects the design.

Next-token prediction

Does this task require probabilistic generation, or does it require precision on specific values?

Classification, summarization, and drafting are probabilistic tasks where the model excels. Extraction of specific authoritative values (account numbers, policy dates, claim amounts) requires verification against the source of truth.

Where design compensates:

  • Generator-verifier loops
  • Code-based evals on extracted values
  • Tool calls to retrieve quantitative data

Knowledge

Does this task depend on information that is rare, contested, recent, or domain-specific in ways that may not be represented in training data?

If yes, the design must bring the knowledge into the context window. Do not rely on the model to supply it.

Where design compensates:

  • Retrieval-augmented generation for stable knowledge
  • Tool calls for live-state data
  • Flagging uncertainty on contested claims

Working memory

Do the inputs fit comfortably in the context window, or does the task require processing inputs that, in aggregate, exceed the window?

Long documents, multi-document tasks, and extended conversations all hit this constraint.

Where design compensates:

  • Chunking strategies
  • Progressive context loading
  • Summarization across turns
  • Pipeline architecture for inputs exceeding the context limit

Steerability

Are the instructions specific, concrete, and verifiable?

Abstract or ambiguous instructions, long reasoning chains, and tasks that require precise numerical or logical computation are all places where the model can drift from intent.

Where design compensates:

  • System prompts with explicit output schemas
  • Structured outputs
  • Code execution for numerical precision
  • Evaluator-optimizer loops

Ground the sizing and value case

Where this would run, and under what rules

These answers decide which products and routes are even available. Leave an answer unknown and Studio keeps it open as a question; it will not assume one for you.

Which rules govern this work? Tick every rule that applies. A ticked rule removes options before any trade-off is weighed.

  • Rules out: Any route where a third party may retain or review the content.
  • Rules out: Any configuration not covered by a signed data agreement for that exact configuration.
  • Rules out: Routes that cannot pin processing to the required region.
  • Rules out: Everything outside the accredited environment.
  • Rules out: Suppliers and products the policy has not approved.
Requirements and evidence

Record what must hold true. A claimed pass or fail requires a source. Simara preserves your statement, but does not independently certify it.

Workload and cost assumptions

Starting numbers are editable planning assumptions, not measurements. Leave unknown optional fields blank. All costs are USD.

Your workload mix

Each row represents a request size. Shares must total 1. Calls count every model step in that request.

Workload 1
Capacity, quotas and context limits

Use measured processing times and account-specific quotas. These estimates cannot establish a response-time percentile or burst capacity.

Who is this report for?

A verified work email starts the report. We send the finished pack to it and keep your brief on record for you.

One report per email address. Revisions of this report do not count as a new one.