Report Manager
ENES

← All articles

Can the AI on your own computer design a report?

· Toni Martir, arquitecto de software

A report does not settle for a passing grade. A letter or a summary can come out so-so and still do their job; a sales report either runs and adds up, or it is no use at all. That makes it a good yardstick for how far artificial intelligence reaches today: there is no room for it to merely look right.

So I ran the test properly. One single brief written as a paragraph, five different models, a score of runs and the same five checks for all of them. Three of the models ran on a home graphics card; two were paid services. I expected the paid ones to win hands down. They did not.

The test

The brief was the same for every model, one only, and far from short. Here it is, whole:

I want you to design a professional sales-by-customer report, grouped by customer: the customer name and code in the header, the customer's total sales in the group footer. The detail must list the items sold to that customer, but only one line per item, that is, the data must be grouped in the SQL by customer and item. At the very end, a bar chart of sales by item (all customers); in each customer's footer, a bar chart of their items too. It must be a professional report; the page header must show the report title (make one up) and Page X of TOTAL.

A report like that marks its own exam. These are the five questions, and not one of them is a matter of taste:

  1. Does it print the total number of pages?
  2. Does it fit the height of the bands to their contents?
  3. Is the list of items there?
  4. Are the amounts formatted?
  5. Does it run first time, with nothing touched?

One brief, five models

Gemma 4 12Bhome card · 271 s
page total: no bands fitted: no runs first time: no
Gemma 4 26Bhome card · 250 s
page total: no bands fitted: no runs first time: yes
Gemma 4 31Bhome card · 672 s
page total: yes bands fitted: yes runs first time: yes
gpt-oss-120Bpaid service · 40 s · three runs
page total: no, all three times bands fitted: no runs first time: no
Gemini 3.8 Flashpaid service · 120 s
page total: yes bands fitted: yes runs first time: yes

The time is for one report built from scratch. The slowest of the five did the best job of it, and it was running on a home graphics card.

It is a tie at the top

Two models passed all five checks: Gemma 4 31B on a home graphics card and Gemini 3.8 Flash in reasoning mode, paid. And that is the news, because I expected only the paid one to fill that box.

The paid one produced the better-looking report, no argument: it put headings on the columns, a subtitle and a units column nobody had asked for. The home one came out plainer, but it got right what decides whether a report is any use — it fitted the height of the bands, twenty pages where others spent forty-seven, and it printed the page total. It took eleven minutes against two.

The surprise is at the bottom. The fastest paid service turned the same brief around in forty seconds, three times running, and all three times the report was worse than the one from the home card: once it left out the whole list of items, which was the body of the brief, and not one of the three printed the page total.

It is worth saying what that «no» in the last column means, because it is no disaster: it means first time. The ones that failed do end up working after a nudge — point the error out to the copilot, or change one expression by hand — and then they print their data and add up. The difference is not whether a report comes out: it is whether it comes out finished. Only two of the five handed over a report with its labels, its column headings and its bands cut to size without anyone touching it, and they are the two at the top. The rest give you something that looks like a report and leave you to finish it.

This is how the best of the five laid it out, exactly as it came from the brief above and with nothing retouched:

The report bands in the designer: a page header with the title in blue and a
subtitle below it; a customer header with the customer name and, under it, the
headings Item Code, Description, Quantity and Total Sales; a detail band one
line tall; a customer footer with its total and a chart; and a report footer
with the grand total and the chart for all customers.

Each grey strip is a band of the report. The «Detail» one — the one that repeats for every item — is a single line tall: that is what separates a twenty-page report from a forty-seven-page one.

How much you ask for at once

That paragraph of a brief looks like one thing, but it carries eight demands inside: group in the SQL, one line per item, the name and code at the top, the total at the bottom, two different charts, the title, the numbering with its total, and the whole thing presentable. And they do not get settled in one sitting: the model asks for the data, describes it, and only then builds the report. It has to reach the end still remembering everything from the start.

That is where the difference lies, and I watched it happen. I asked the twelve billion model to fix one specific fault, just one, and it fixed it — including turning on the two-pass setting, which is what makes the page total print, and which nobody had mentioned to it. That same model, given the whole brief, had left out the list of items.

It knows the rule; what it lacks is the stamina to hold on to it all the way. In another run it wrote the rule itself in its own reasoning — «this needs the two-pass setting, I will do it in the next step» — and by the time the next step came it no longer remembered.

Hence the practical advice, which is also how anyone works with the designer: build the report in pieces and a small model gets much further. First the data. Then the header. Then the detail. Then the totals. Each request is small, you see the result and you correct it before moving on. With a big model you can afford the whole paragraph at once; with a small one, the whole paragraph at once is exactly what you should not ask for.

So, can it?

Yes, and the condition is not the one I expected. If you want the whole report from a paragraph, you need a model that fits in some twenty-four gigabytes of card memory. With twelve there is another way: build it in pieces. The small model gets along well that way, and it writes at three times the speed of the winner.

And size is deceptive. The one that came off worst of the two paid services is the largest of them all on paper — one hundred and sixteen billion parameters — but it only puts five billion of them to work on each word it writes, and it belongs to an earlier generation. The one that tied at the top from a home card has thirty-one billion and uses all of them, and it is from this year. Mid-sized and modern beats large and dated.

The frontier of local AI is not where I thought. It is not the money, nor the speed, nor even the size: it is the relation between how much you ask for at once and how much the model has in it.

In Report Manager the copilot that designs the report is optional, and when you do use it the work can be done by a model running on the customer's own machine, served by the Reportman Agent: no monthly fee, with the customer providing the graphics card. The report data never leaves the company, which is the other side of the same question and one I covered separately in how much is your data privacy worth?. The designer and the server are free, in the sense that there is nothing to pay to use them.

Toni Martir, software architect — 9 October 2026

The test setup

Sources

Toni Martir, arquitecto de software · · sources checked on