Марія БровінськаAI Eng
24 July 2026, 09:01
2026-07-24
A man from Aitivtsev arranged an exam for four LLM students on real data. Here's what happened
Product Marketing Manager at Railsware Artem Sagaydak organized an exam for four LLMs on real data. He shared this experience with dev.ua. Below is Artem’s direct speech.
Product Marketing Manager at Railsware Artem Sagaydak organized an exam for four LLMs on real data. He shared this experience with dev.ua. Below is Artem’s direct speech.
The essence of the experiment
You can buy the most expensive subscription to your favorite AI, write a prompt perfectly, and still get an answer that doesn’t match reality. Your cortisol levels are rising, the time before the meeting where you have to present data is running out, and you start thinking about all those LinkedIn influencers with advice and «100% working» prompts.
Like all marketers, I use AI for many tasks. Often working with them is more exhausting than it gives a really cool result. So I decided to conduct a study, namely to test the four most popular models GPT-5.5, Claude Opus 4.8, Gemini 3.1 Pro and Grok 4.3 and compare the results.
I loaded the same datasets through Coupler.io: funnel, ICP breakdowns, churn analysis, campaign performance, etc. In Agent (Medium) mode in Cursor, I asked ten questions in the same order: find data gaps, build customer profiles, analyze the funnel, assess churn risk, calculate lost revenue, and offer strategic recommendations.
To make the environment more realistic, I fed the models data that sometimes lacked clarity: they contradicted each other or did not allow for a confident conclusion. This helped me better understand the behavior of each model and understand when to use it.
So if you’re interested in seeing the results, here’s what I came up with.
What are the conclusions based on?
I could, of course, share the results of all ten questions. But in order not to overwhelm you, I will focus on the three most interesting ones (in fact, it was difficult to choose). They are the ones that best show the differences in the behavior of models in popular cases.
Test 1. Does the model recognize that there is no data?
Question: Let’s analyze Q4 2025. Which campaigns have the highest CPA relative to their trial-to-paid conversion? Where should we cut spend?
Here, we tested whether the model could detect a gap in the data in the middle of a task and still produce something useful: work with what was available, explain what was missing, and suggest a specific next step.
What we received:
GPT-5.5
I noticed a lack of data and built a ranked action table with color-coded labels (Cut / Cap / Tighten / Protect).
Minus: the recommendations were based on conversion rates, which he himself called unreliable, without noticing the contradiction.
Gemini 3.1 Pro
Confirmed the data gap, found the right dataset, and asked which advertising platform to connect to.
Minus: did not offer any interim analysis based on already available information.
Grok 4.3
I quickly noticed that there was no data on expenses.
I ranked the segments by conversion to payment (e-commerce 8%, finance 4,5%) and advised to reduce «Finance» and «Analysts» first.
Concise and useful, but without visualization.
Claude Opus 4.8
I didn’t see it as a data problem: I found costs at the individual campaign level and answered the question right away.
I identified a group of campaigns targeting competitors (52% of Q4 budget, no paid conversions) and suggested cutting them.
advised to save budget for Brand Search and drew attention to possible attribution artifacts in Meta and LinkedIn data.
Test 3. Can it work with combinations of sources?
Question: What are the most common data source combinations our e-commerce customers connect and which combos correlate with the highest paid conversion rate?
I tested whether the models were able to detect correlations between connector combinations and paid conversions, and whether they could distinguish between volume (which is the most common) and efficiency (which results in conversions).
What we received:
GPT-5.5
Built a full panel (canvas).
Key finding: Paid conversion increases with the number of sources. One gives 6,3%, three — 31,4%, five or more — 48,2%.
Best consistent combination: Facebook Ads + Google Ads + Google Analytics + Shopify = 45,5% (11 accounts); most common single source: Shopify (238 accounts).
Gemini 3.1 Pro
Showed the same data regarding combinations.
Added analysis of the «source → destination» chain, which was not in GPT: Google Ads → BigQuery = 68,4%, Facebook Ads → BigQuery = 60,9%, Klaviyo → Claude = 53,8%.
Conclusion: where you send data is just as important as what you connect.
Grok 4.3
I wrote down the most common sources: Facebook Ads (545), Shopify (508), Google Ads (380).
Added the connection Shopify + Facebook Ads + Google Ads — about 49,6% conversion.
Emphasized: settings with multiple data sources signal intent 7–8 times stronger than a single source. Pure analysis, but no visualization.
Claude Opus 4.8
Responded in text, narrowing the sample to e-commerce organizations Q4 2025.
Paid conversion increases about 5x from one source to three or more; Shopify is the anchor, Facebook Ads alone is a dead end.
The biggest inhibitor: most e-com registrations don’t connect any sources at all.
Test 9. Can it synthesize the entire dataset and create an effective recommendation?
Question: If you were our Product Marketing Manager, what one thing would you change based on this data?
Here, it was interesting to see whether the model was able to synthesize the entire dataset into one well-founded strategic recommendation and present it as a PMM solution, rather than as an analyst’s observation.
What we received:
GPT-5.5
He proposed making e-commerce the main focus of the PMM strategy and rebuilding the top of the funnel around e-commerce scenarios.
Warning: trial period data is unreliable (the number of paid exceeds trial), so don’t rely too much on it.
Concise and correct, but without a specific plan of action.
Gemini 3.1 Pro
He advised shifting positioning, communication, and marketing budget toward e-commerce.
Argument: The company is attracting the wrong audience. Agencies and analysts have low conversion and high churn, and the most valuable segment receives the least attention.
Specific steps: change the message on the main page, redistribute the advertising budget, focus content marketing on e-commerce scenarios.
Grok 4.3
The priority should be on engagement and communication for e-commerce; expenses for analysts and finance should be reduced or more strictly qualified.
He explained the strategic logic in the most detailed way: the company pays to attract users for whom the product does not yet create sufficient value.
He noted that this will have a greater impact than further optimizing onboarding.
Claude Opus 4.8
I rethought the answer: I shifted the focus from the segment to the wedge — the main point of entry into the market.
He suggested abandoning the universal «connect any data» and rebuilding the top of the funnel and the first experience around business owners and corporate email marketers who start with a ready-made dashboard template.
I supported it with numbers: the difference in conversion between corporate and public emails, high conversion of template users, and an assessment of the potential effect.
Which LLM model should I choose for data analysis?
After ten tests, the answer was not as obvious as we would like. All models can analyze data and can be good helpers in the work. However, they do it very differently.
Some are faster at finding patterns, others are better at explaining their conclusions, and still others are more confident in moving from analysis to specific recommendations. Therefore, the choice of model depends not so much on which one is the best, but on what result you want to get.
I also add a brief summary of the behavior of LLM models.
How to increase the chances of better cooperation with LLM models?
In short, treat the model not as a magic bullet, but as a very fast junior colleague. It can find a pattern in a minute that you would have spent an hour looking for. But it can also confidently go the wrong way if given a vague task or questionable data.
Here are some tips that really improve the result:
Don’t stop at the first answer. You can find the most interesting insights if you start to expand the model with questions: «Why?», «What influenced this?», «What else could this mean?».
First, check the data, then ask it to analyze it. If the dataset is incomplete or contains contradictions, the model will not always stop on its own. It will most likely try to find the answer anyway.
Be specific in your questions. «What’s going on?» rarely leads to useful conclusions. «Why don’t 68% of users get past the first step of onboarding?» is a completely different conversation.
Don’t forget about context. If you have your own definitions of metrics, segments, or business terms, provide them to the model. Otherwise, it will fill in the gaps with its own assumptions.
Don’t just read the pretty graphs. The most valuable warnings and explanations are often hidden in the text response.
And most importantly, don’t delegate the last word to AI. The model is great at finding patterns, but decisions that affect the business should still be checked against common sense and verified with the original source.
Models have learned to respond, but not to doubt yet
Ten tests showed a strange thing: the models hardly argue with each other about the answers. They argue only about the presentation. Where the data is clean, the four different architectures come to the same conclusion. And that’s good news for anyone who was afraid to trust the numbers in the chat.
The bad news is hidden elsewhere. None of the four asked whether the question was even correctly posed. Did not suggest looking in the CRM, where there is already a ready-made list of risky clients. Did not calculate how much the strategy reversal, which she herself advises, would cost. The models play brilliantly, but only within the boundaries of the board that was drawn for them.
Therefore, the division of labor for the near future looks like this: the model searches and counts, the person puts a frame and doubts. The first already costs a penny and takes minutes. The second, as it turned out, is not yet sold for any token. Perhaps new models will change this balance, but that is a topic for another article.