In its 30 September 2026 analysis, Artificial Analysis reports an Intelligence Index score of 53 for Gemini 4 Argon (high), matching GPT-6 Astra (max). Cost per Task is $1.99 with the introductory discount and $3.98 at standard prices. Access is limited to selected users, and the promotion has no confirmed end date. Source: Artificial Analysis.

For a team choosing a model for an agent, that gives Argon a place in the test queue once access is available. We would base a deployment decision on the cost of completing the actual work, including corrections, retries and review. What follows is our proposed evaluation approach. We have not run our own measurements of this model.

Intelligence Index Cost per Task: Gemini 4 Argon (high), $1.99 at introductory prices and $3.98 at standard prices.
Chart: Lazarus Systems · data: Artificial Analysis, 30 September 2026 ↗

What the $1.99 figure measures

In the Artificial Analysis methodology, Cost per Task is the weighted average cost of a task in the Intelligence Index suite. It accounts for input, cached and output token usage at the applicable prices. The choice of tasks and benchmark weights affects the result.

That figure does not price an invoice check or a contract review. Our agent might need several model calls, a search request and a retry after a failed step. Some answers will need a person to check them. Each of those costs belongs in the estimate.

We would budget a pilot at standard rates. A promotion can reduce the bill during an experiment; the project should remain affordable after it ends. If the financial case depends entirely on the discount, the deployment decision should say so.

A trial using real documents

Suppose an agent needs to compare purchase orders with invoices and flag discrepancies. We would prepare a fixed set of documents with expected outcomes. It should include poor scans, missing pages, different currencies and documents that cannot be matched to an order with confidence.

The acceptance criteria need to be written before testing starts. Producing an answer is insufficient. The agent should select the right documents, read the amounts correctly and identify cases that need an operator's decision. Where possible, we would hide the model's name from the reviewer so that expectations about the provider do not influence the assessment.

Each model would receive the same tools and constraints. We would record the time taken for the full run, the cost of every call, the number of retries and the time needed to review the answer. False acceptances deserve a separate count: cases where the agent marks a document as correct despite a discrepancy.

A close-up of a classical bust against black.
Editorial illustration.

Conditions for switching models

The acceptance threshold should follow the business process. An agent sorting documents can tolerate a different error rate from one authorising payments. A useful comparison starts by deciding which actions the agent can take independently and which require approval.

The test record should include the exact model version, reasoning settings, price schedule and measurement date. We would attach examples of failures and the cost per successfully completed case to the decision. That record gives the team a basis for checking whether a later model version improves the results.

Benchmark figures come from the Artificial Analysis publication linked above. The proposed test and deployment recommendations are commentary by Lazarus Systems.

LAZARUS SYSTEMSLet’s talk about your project ↗