Groovy Web's 2026 hiring research, cited by Abbacus Technologies, puts the average cost of hiring an AI engineer at upwards of 185,000 US dollars annually with a hiring cycle running four to six months, which is exactly why a mis-hire caught early is far cheaper than one discovered after a failed launch. Amplence's 2026 vendor guide found that three or more of its twelve documented red flags appearing together in a single vendor is a near-certain predictor of project overrun or failure, drawn from patterns observed across the AI services market in 2025 and 2026. RocketFarm Studios' 2026 buyer's guide is direct about where these signs actually surface: most AI products fail after the demo, once real data, security rules, and performance requirements enter the picture and the polished prototype falls apart.
The twelve signs below cover the portfolio, the conversation, and the contract, the three places a red flag reliably shows up before you have committed a budget to a developer who cannot actually deliver. Hire an AI developer who clears all twelve, and you have filtered out most of the risk before the first invoice goes out.
1. The Portfolio Is All Jupyter Notebooks, No Production Deployments
Interexy's 2026 hiring guide is blunt about this pattern: a portfolio full of Jupyter notebooks is great for exploration and terrible evidence of production ability, since notebooks skip packaging, testing, and deployment entirely. Automely's 2026 hiring research frames the filter that matters: has the candidate shipped an AI system that real users interacted with, at real scale, with real consequences for failure, not a proof of concept or a demo that only works on carefully selected test data. If every project in a portfolio is a notebook with a impressive-sounding description and no link to a live system, that absence is the answer.
2. Claims Expertise Across Every AI Specialisation at Once
Interexy's guide calls this out directly: a candidate claiming expertise in LLMs, computer vision, NLP, reinforcement learning, and quantum computing all at once is not being efficient with their resume, they are exaggerating. TechExactly's 2026 hiring guide echoes the same warning for candidates who claim mastery across LLMs, computer vision, robotics, and blockchain simultaneously. Genuine specialists tend to be precise about the boundaries of their expertise, not expansive about them, which is worth probing directly with the kind of structured ML interview questions that separate real depth from a broad but shallow resume.
3. No GitHub or Private Portfolio, Just "I Signed an NDA"
An NDA is sometimes the real explanation for a thin public portfolio, but Interexy's research notes it is usually an excuse rather than the full story, and a candidate with genuine experience can almost always produce a personal project or an anonymised code sample even when client work is confidential. The AI agent hiring red flags post covers this pattern in more depth specifically for agent development freelancers, where the same excuse shows up often given how recent most production agent work actually is.
4. Cannot Describe a Production System That Actually Broke
Automely's 2026 interview research recommends asking directly: walk me through a production AI agent that broke, what happened. Real developers have specific, detailed answers involving memory persistence issues, vector database retrieval under real query volumes, or a failed tool call mid-task; tutorial-level candidates give vague generalities about "working with LLMs" instead. DestiLabs' 2026 hiring guide frames the same test differently: ask to see something shipped to production, not a demo, since this single question filters out most of the field before a formal interview even starts. A developer proposing custom AI agent solutions who has never had one fail in production has probably never run one at real scale.
5. "The Model Just Works" Is the Whole Explanation
TechExactly's 2026 guide flags this specifically: if a candidate says the model just works but cannot elaborate on data quality, evaluation methodology, or monitoring when asked, that vagueness is the problem, not a sign of confidence. A developer who has actually built and maintained a production model can talk through exactly how they know it works and what would tell them if it stopped.
6. Overpromises Specific Results Before Seeing Your Data
Louis Innovations' 2026 hiring research calls overpromising the most dangerous red flag in AI hiring, since AI development involves genuine uncertainty and honest developers acknowledge it rather than guaranteeing an outcome before understanding the data. RocketFarm Studios' buyer's guide lists this as a distinct red flag in its own right: any vendor promising a specific timeline without having seen your data first is guessing, not scoping.
7. Case Studies Read Like Marketing Copy, Not Outcomes
RocketFarm Studios draws a sharp line here: "built a chatbot" tells you nothing useful, while a case study describing a claims-processing AI agent that reduced manual review time by 62 percent across 14,000 monthly transactions tells you everything you need to evaluate the work. A portfolio full of the first kind of description and none of the second is a portfolio that has not been asked to prove anything yet.
8. No Transparent Pricing or Defined Acceptance Criteria
Amplence's 2026 vendor research lists no transparent pricing and vague deliverables with no acceptance criteria among its most common and most costly red flags, alongside multi-month "AI strategy" engagements that produce no working code. A developer who cannot describe how the project will be measured as complete before the contract is signed is setting up a scope dispute for later in the engagement rather than avoiding one now.
9. No Data Privacy Plan or Explainability Discussion
Amplence's guide adds two compliance-focused red flags that matter most for regulated industries: no data privacy plan and no explainability discussion, citing FTC Operation AI Comply enforcement activity and EU AI Act Articles 13 and 14 as the regulatory backdrop making these non-negotiable rather than nice-to-have. A developer who has not thought through how a model's decisions will be explained to an end user or a regulator has skipped a requirement that surfaces expensively later, not a theoretical concern.
10. Locked Into a Single LLM Provider With No Rationale
Amplence's research names only using one LLM as a specific red flag, since a developer who defaults to a single provider without discussing trade-offs is either inexperienced with the alternatives or building in a dependency you did not ask for. A developer proposing generative AI development services should be able to explain why a specific model fits your use case and cost profile, not simply which one they personally prefer to use.
11. No Plan for What Happens After Launch
DestiLabs' 2026 hiring guide asks the question directly: what happens after launch, since monitoring, alerting, and tuning matter because AI systems drift as data and usage change, and a candidate who only talks about the build has never actually run a system long enough to see that drift happen. Amplence lists no post-launch support among its twelve red flags for the same reason: a project that ends at deployment is not actually finished.
12. Asks for Payment or Equity Before Discussing the Problem
Interexy's guide is direct here: equity is for co-founders, not for a first conversation, and a candidate who asks for it before proving anything is prioritising their own upside over your problem. Louis Innovations adds a related pattern worth watching for in the same conversation: resume buzzwords without depth and an unwillingness to discuss past failures openly, since a developer unwilling to talk through what went wrong on a previous project is unlikely to tell you when something goes wrong on yours.
12 Red Flags at a Glance
|
Red Flag |
What Good Looks Like Instead |
|---|---|
|
Notebook-only portfolio |
Links to live, production systems real users interact with |
|
Claims expertise in everything |
Clear, specific boundaries around their actual specialisation |
|
No GitHub, only "NDA" |
A personal project or anonymised code sample on request |
|
Can't describe a production failure |
A specific, detailed story about a system that broke and how it was fixed |
|
"The model just works" |
A clear explanation of data quality, evaluation, and monitoring |
|
Overpromises before seeing your data |
Acknowledges uncertainty and asks about your data first |
|
Marketing-copy case studies |
Case studies with named, measurable outcomes |
|
No transparent pricing or acceptance criteria |
A defined scope and a clear definition of done |
|
No data privacy or explainability plan |
A documented plan for both, especially in regulated industries |
|
Locked into one LLM with no rationale |
A clear explanation of why a specific model fits your use case |
|
No post-launch plan |
A defined monitoring, alerting, and retraining plan |
|
Asks for equity before proving value |
Asks about the problem first, discusses terms once value is proven |
Twelve Flags, One Filter
None of these twelve signs is disqualifying on its own in every case, but they cluster for a reason: developers who have actually shipped production AI systems tend to clear all twelve without hesitation, because they have lived through the consequences of skipping each one. The Gen AI portfolio red flags post covers a narrower version of this checklist specifically for generative AI and LLM-focused candidates.
Hire AI and ML developers who can answer every one of these twelve questions with specifics rather than reassurance, and treat that clarity as the actual qualification, not the polish of the pitch.
