
Match BERT With 10% Labels: Medium LLM Intent Detection, EU Ready
Match BERT With 10% Labels: Medium LLM Intent Detection, EU Ready

The practical best practice for modern NLU intent detection is to use joint intent+slot architectures or fine-tuned medium-scale LLMs that deliver competitive performance with far less annotated data. The main trade-offs sit between data and compute budgets, between fixed-threshold and threshold-free multi-intent handling, and between model capability and GDPR-driven data minimization. The sections below walk through architecture choices, evaluation metrics, recent research, compliance constraints, and a concrete implementation path.
TL;DR:
- Fine-tuning medium-scale LLMs with only 10% of annotated data can match or outperform fully trained joint models, reducing data needs significantly.
- Threshold-free multi-intent detection predicts the number of intents in an utterance, improving accuracy over fixed confidence thresholds.
- GDPR compliance requires early data protection impact assessments, anonymization, and strict retention policies before training models.
- Using joint intent and slot models enhances accuracy and information sharing, especially with transformer encoders and adapter-based fine-tuning.
- Effective deployment involves continuous monitoring, retraining, and deliberate fallback design to handle confidence drift and evolving traffic patterns.
Table of Contents
- Where Intent Detection Fits in the NLU Stack
- Architectures: Pipelines, Joint Models, and Medium-Scale LLM Fine-Tuning
- Benchmarks, Metrics, and Confidence Thresholds
- Advances Worth Adopting Now
- Data Protection Requirements for NLU Agents in Europe
- A Practical Roadmap from Data to Deployment
- How We Approach This in Production Projects
- Why We’re a Practical Starting Point for Your NLU Project
- FAQ
- Sources
Where Intent Detection Fits in the NLU Stack
Intent detection classifies what a user wants at the utterance level: book a flight, cancel an order, escalate to a human. Slot filling works at the token level, pulling out the specific values, a city name, a date, an order number, that the intent needs to act on. Both tasks typically sit inside a larger pipeline: speech recognition feeds raw text to the NLU layer, which hands structured output to dialogue state tracking, then to a policy engine, then to response generation.
Single-intent detection covers most transactional requests (“reset my password”). Multi-intent detection matters when users bundle requests in one sentence, a common pattern in voice assistants and WhatsApp-style messaging threads.
- Intent detection answers “what does the user want,” slot filling answers “with which details.”
- A typical pipeline runs ASR, then NLU, then dialogue state tracking, then a policy layer, then natural language generation.
- Multi-intent detection becomes necessary once users naturally chain requests in a single message.
Architectures: Pipelines, Joint Models, and Medium-Scale LLM Fine-Tuning
Early systems treated intent classification and slot filling as separate problems, often using CRFs or sequence taggers trained independently. That separation loses information: knowing the slot “Paris” was just filled tells you something about whether the intent is “book flight” or “check weather,” and a pipeline architecture cannot pass that signal backward.
Joint models fix this by sharing an encoder between both tasks and adding task-specific heads, sometimes with gating mechanisms that let slot predictions inform intent predictions and vice versa. A review of joint learning approaches traces this from classical statistical methods through deep learning, finding that joint architectures consistently beat separate pipelines on both intent accuracy and slot F1 because the two tasks share mutually reinforcing signal.
Transformer encoders (BERT-style models with intent and slot heads attached) became the default for several years. The newer shift is fine-tuning medium-scale LLMs, roughly 8 billion parameters, for the same joint task. A 2025 COLING industry paper found that a fine-tuned medium-scale LLM using only 10% of the annotated training data matched or beat a fully trained JointBERT baseline, using LoRA adapters and 8-bit quantization to keep hardware requirements inside the reach of a single mid-range GPU.
- CRF and pipeline sequence taggers were the baseline before deep joint models arrived.
- Joint encoders share representations across intent and slot heads, improving both tasks simultaneously.
- LoRA plus 8-bit quantization lets medium-scale LLM fine-tuning run on 12 to 48 GB GPUs for training and as little as 8 GB for inference.
- On-premise deployment becomes realistic at this scale, which matters for teams with data residency constraints.
Pro Tip: Start LoRA fine-tuning with a frozen base model and train only the adapter layers on your joint intent and slot-labeled data; quantize to 8-bit only after validating accuracy on full precision.
Benchmarks, Metrics, and Confidence Thresholds
Benchmarking intent detection means picking datasets that match your domain’s complexity. ATIS covers single-domain airline queries with simple slots. SNIPS spans several consumer domains. MASSIVE adds large-scale multilingual coverage. NLU++ targets fine-grained, overlapping intents closer to real customer service traffic.
| Dataset | Domain scope | What it reveals |
|---|---|---|
| ATIS | Single domain, airline | Baseline performance on simple slot structures |
| SNIPS | Multiple consumer domains | Cross-domain generalization |
| MASSIVE | Multilingual, broad domains | Performance across languages |
| NLU++ | Fine-grained, overlapping intents | Behavior on ambiguous, real-world phrasing |
Track intent accuracy, slot F1, and joint exact match (both predictions correct together) alongside latency and memory footprint, since a model that wins on accuracy but misses a latency budget is not deployable.
- Confidence scores typically range from 0.0 to 1.0 and drive fallback triggers in production systems.
- Raising a confidence threshold increases precision but also increases how often the system falls back to a human or a clarifying question.
- Monitor confidence distributions over time, not just at launch, since drift changes what “confident” means for your traffic.
Advances Worth Adopting Now
Fixed confidence thresholds are brittle: a single cutoff rarely fits every intent or every traffic pattern. Threshold-free multi-intent models solve this with an auxiliary Intent Number Detection task that predicts how many intents are present in an utterance, then selects that many top-ranked labels instead of relying on a probability cutoff. This approach has shown higher accuracy on multi-intent benchmarks than threshold-based baselines.
Adding sentiment or emotion as an auxiliary signal can help disambiguate intent in some domains, though the research on sentiment-aided intent detection shows the benefit varies by domain and needs careful feature engineering rather than a drop-in addition.
- Intent Number Detection replaces a fixed threshold with a predicted count, then selects that many top labels.
- Sentiment and emotion features can reduce misclassification but require domain-specific validation before relying on them.
- Medium-scale LLM fine-tuning reduces the annotation burden needed for cross-lingual transfer, since base models already carry multilingual representations.
Data Protection Requirements for NLU Agents in Europe
NLU systems that process customer utterances handle personal data, which puts them squarely inside GDPR’s transparency, data minimization, and purpose limitation requirements. Article 9 adds stricter rules when conversations touch special categories of data, health mentions or similar sensitive topics.
The CNIL recommends running a Data Protection Impact Assessment (AIPD) at the design stage for agentic AI systems, before training begins, and specifically warns against coupling raw transcripts with identifying metadata. Our CNIL checklist for GDPR and AI walks through timing the AIPD correctly.
- Pseudonymize or hash user identifiers before they touch training logs or analytics dashboards.
- Set retention policies that delete raw transcripts on a schedule rather than keeping them indefinitely by default.
- Offer clear opt-outs and consent patterns before conversations are used for model improvement.
Pro Tip: Run the AIPD before the first training run, not after a model is already in production, since retrofitting privacy controls into a trained pipeline is far costlier than designing them in. The EU AI Act’s risk-based obligations, including requirements for higher-risk AI systems, are phasing in on a schedule that makes early compliance planning worthwhile.
A Practical Roadmap from Data to Deployment
Getting from a raw dataset to a reliable production intent classifier follows a predictable sequence, though teams skip steps under deadline pressure more often than they should.
- Build an annotation strategy that prioritizes intent diversity and edge cases over raw volume, using active learning to flag the utterances most likely to improve the model.
- Choose fine-tuning a medium-scale LLM when annotated data is scarce and you have access to a GPU with at least 12 GB VRAM for LoRA training; choose a lightweight joint transformer model when data is abundant and inference latency is tight.
- Instrument monitoring before launch: confusion matrices by intent pair, confidence distribution dashboards, and drift alerts tied to a defined retrain cadence.
- Design fallback UX deliberately, with confirmation prompts for low-confidence predictions, graceful degradation rather than silent failure, and a clear human-in-the-loop escalation path.
Our guide to building an AI agent from scratch covers the engineering patterns behind step one and two in more depth, and our overview of common chatbot mistakes catalogs the fallback UX failures that show up most often in live support deployments.
Pro Tip: Treat your confusion matrix as a living document: review it after every retrain, not just at initial launch, since new intents and phrasing drift in faster than most teams expect.
How We Approach This in Production Projects
We treat the research above as a starting point, not an end state. Our projects typically move from a short audit, to a fast proof of concept built around a joint model or a fine-tuned medium-scale LLM depending on available data, to a production phase where monitoring and GDPR-aware deployment controls are built in from day one rather than bolted on later. Our governance guide for AI decision-makers outlines how we structure accountability across that flow.

Why We’re a Practical Starting Point for Your NLU Project
Reading the research is one step. Building a joint model, fine-tuning a medium-scale LLM on your own support tickets, and wiring in GDPR-compliant logging is a different kind of work, and it is the work we do daily for clients moving from a research idea to a live chatbot. We offer custom chatbot development, AI integration consulting, and tailored automation built around your actual data, not a generic template.

If you already run a Shopify store, our Shopify app development work includes the Plan Starter at $19 per month and the Plan Pro at $49 per month for ongoing chatbot and automation needs. For broader integration work, from WhatsApp Business deployments to custom n8n and Make automations, our AI solutions overview lays out the full range of services, and our Aria by BotiqueAI product shows what a production-grade customer assistant looks like once the underlying intent detection is solid. Reach out for a project audit and we will tell you plainly whether a joint model or a medium-scale LLM fits your data better before any contract is signed.
— Botiqueai
FAQ
What is the difference between intent detection and slot filling?
Intent detection identifies the overall goal of a user’s message, such as booking a flight, while slot filling extracts the specific details within that message, like the destination city or travel date. Modern systems typically handle both with a joint model that shares information between the two tasks rather than running them as separate steps.
How does threshold-free multi-intent detection work?
Instead of relying on a fixed confidence cutoff to decide how many intents are present, a threshold-free approach trains an auxiliary task to predict the number of intents in an utterance, then selects that many top-ranked labels. This avoids the brittleness of a single manually tuned threshold across varied traffic.
Can fine-tuning a medium-scale LLM reduce how much training data I need?
Yes. A 2025 industry paper found that fine-tuning a medium-scale LLM with LoRA and 8-bit quantization matched or exceeded a fully trained BERT-based joint model while using only 10% of the annotated data. This makes the approach particularly useful for teams without large labeled datasets.
What does the CNIL recommend for AI systems handling conversational data?
The CNIL recommends conducting a Data Protection Impact Assessment at the design stage, before training begins, and advises against storing raw transcripts alongside identifying metadata. Pseudonymization and clear retention policies are core practical controls.
Does BotiqueAI build custom intent detection systems for businesses?
We develop custom chatbots and NLU integrations tailored to a client’s own data and compliance requirements, including GDPR-aware deployment. Pricing for our packaged Shopify chatbot app starts with the Plan Starter at $19 per month, while custom consulting and integration projects are scoped individually.
Sources
- A joint learning classification for intent detection and slot filling from classical to deep learning: a review
- AI system development: CNIL’s recommendations to comply with the GDPR
- Threshold-free multi-intent NLU model (TFMN) and Intent Number Detection