Choosing an AI automation partner in Spain is a security and operations decision, not a software purchase. These are the criteria you can verify before you sign, and the red flags worth a second look.
TL;DR: Choose on criteria you can verify rather than on the demo: a security posture built on least privilege and the four production controls, working knowledge of GDPR and the EU AI Act, evidence instead of slides, a scoped and reversible pilot before scaling, and a partner who works in your languages with clear control of your data.
In practice, what decides the outcome of an AI project is not the model but the scope, the permissions, and who answers when the agent gets something wrong. Choosing a partner decides who will touch your data, who defines what an agent can reach, and who takes responsibility when something breaks in production. That is a security and operations decision before it is a software purchase, and it deserves the same scrutiny you would give any new access to your systems.
What is really on the table: data access, permission boundaries, and accountability when it fails.
Good sign: you can change partner without rebuilding the operation.
Bad sign: the conversation starts and ends with the demo.
Criterion 1: security posture, not just capability
Ask how they limit what the agent can reach. The principle is least privilege: exactly the tools and data the task requires, no more and no less, and only for as long as it is needed. Then ask which controls will be live before anything reaches a customer. A partner who cannot describe what happens when the agent gets it wrong does not yet have a production-ready system, however well the demo runs.
Granular rollback: undo one action without tearing down the whole workflow.
A human review queue for high-risk actions.
Searchable, traceable logs with correlation identifiers.
Trusted evals with quantified thresholds before autonomy widens.
Criterion 2: GDPR and the EU AI Act in the design
A partner working in Spain should be able to say plainly what personal data each workflow touches, on what legal basis, where it is processed, how long it is kept, and which third parties are involved. The EU AI Act adds a second question: which risk category the use case falls into, and what transparency, human-oversight, and documentation obligations follow from it. Neither is a badge you buy: both are design decisions taken while mapping the workflow, before anything is built.
What personal data each workflow touches, and on what legal basis.
Where it is processed, how long it is kept, and which providers are involved.
What transparency and human oversight the use case requires.
Criterion 3: evidence instead of demos
A demo shows the happy path. Ask for what it does not show: how the system is evaluated and against which thresholds, what the logs look like and who can search them, what the rollback plan is, and what documentation you keep if you change partner tomorrow. Ask for a real failure and what changed afterwards — that answer separates people who have operated systems from people who have only presented them. Slides describe capability; a written scope, an eval set, and a recovery path are evidence.
Ask for an eval set with thresholds, not a hand-picked demo.
Ask for a real failure and what was fixed after it.
Ask what documentation and access stay with you at the end.
Criterion 4: start small, reversible, and measurable
A good partner pushes you toward one narrow first workflow rather than a full transformation. Scoped: one workflow, with defined inputs and outputs. Reversible: nothing irreversible leaves without human approval. Measurable: success criteria written before the build, so at the end you can decide on evidence whether to widen it, fix it, or stop. Anyone proposing to automate the whole operation at once is asking you to carry the risk of them learning on you.
Criterion 5: language, proximity, and control of your data
A lot of operational work in Spain happens in more than one language: a hotel answering in Spanish, Catalan, and English on the same day. A partner who works in your languages can read your real inputs — customer emails, listings, tickets — instead of a translated sample, and can train the team in the language it actually works in. Proximity matters for the same reason: sessions with the people doing the work, in their working hours. And ask about deployment options: whether the workflow can run with clear data boundaries, or as a private deployment when the data justifies it.
Works in the languages of your operation: Spanish, Catalan, English.
Sits with the team doing the work, not only with management.
Offers clear data boundaries and private deployment when it is needed.
How to compare proposals, and the red flags to watch
Put proposals on the same axes rather than comparing totals: scope in writing, who does the work, which controls ship with it, what happens after launch, and what you own at the end. On price, what matters is transparency rather than the number: you should be able to see what is fixed and what is variable, what happens to the budget if scope changes, whether model or infrastructure usage is billed separately, and what you stop paying if you stop. A proposal that cannot be broken down that way cannot be compared either.
Vague scope: "we automate your processes with AI", with no workflow and no concrete deliverable.
No security story: nobody explains permissions, logs, or how an error gets reversed.
Big-bang project: the whole operation in a single phase, with no pilot first.
Lock-in: no documentation, no access, and no way to continue without them.
Opaque pricing: fixed and variable are not separated, and scope changes are undefined.
Where Paput fits
Paput is one of the options that meet these criteria, and this section deserves the same skepticism as the rest. It is a small studio based in Menorca, working remotely across Spain and Europe in Spanish, Catalan, and English. It designs custom agents and hardens existing AI workflows with the same four production controls on every build, least privilege on tool and data access, and human approval on sensitive actions. The entry point is an AI and security audit that maps automatable workflows, access, data boundaries, and next steps, followed by a scoped pilot before anything scales. If, once you compare, another option fits your case better, that is the right answer.
Questions buyers ask
What should I ask for before signing?
The scope of one workflow in writing, what data it touches and on what legal basis, which controls ship with it — rollback, human review, logs, and evals — measurable success criteria, and what documentation and access stay with you at the end. If any of that is missing from the proposal, ask for it before signing, not after.
How do I know a partner understands agentic security?
Ask them to describe the worst failure the workflow could produce and how it is contained: what the agent can reach, which actions need approval, what gets logged, and how it is reversed. A generic answer about "secure models" means there is no design behind it; a specific answer can be audited.
Does the partner have to be local?
Not necessarily, but it helps when the operation runs in several languages or the team needs in-person support. What is not negotiable is that someone understands your data, the language your team works in, and the European regulatory framework that applies to you.
How do I compare prices when every proposal has a different format?
Do not compare totals: compare what is included. Ask them to separate the fixed part from the variable one, what happens if scope changes, whether model and infrastructure usage is billed separately, and what you keep paying after launch. A proposal that cannot be broken down that way cannot be compared either.
AI operator field notes
illmethinks.io publishes source-transparent notes on AI agents, tools, and operational risk monitored by Paput.ai.