Capability evidence
Assess the language and evidence decisions hidden behind the model output.
A useful assessment includes multilingual and mixed-script text, disputed annotations, unseen terminology, sparse relevance judgments, a strong simple baseline, a fluent unsupported answer, an adversarial instruction, and a delayed correction. It should reveal whether the person can preserve text and context, choose task-specific evidence, expose uncertainty, and route failure to the right human owner.
01Language, text, corpus, and annotation boundary
Give the specialist composed and decomposed characters, combining marks, emoji sequences, mixed left-to-right and right-to-left text, language and script ambiguity, code switching, markup, OCR noise, changing documents, sensitive text, reused passages, span labels, and annotator disagreement. Ask for storage, segmentation, metadata, corpus, annotation, and correction design.
Confirm: The person preserves original text and stable document identity, distinguishes code point, grapheme, word, token and linguistic unit, treats normalization as a versioned transformation, carries language and direction metadata, protects offsets through edits, records source and rights, measures corpus coverage, designs clear guidelines, preserves disagreement, and avoids presenting heuristic language detection as fact.
02Search, retrieval, classification, and extraction
Use ambiguous and rare queries, synonyms, negation, typos, multiple languages, unseen terms, long documents, duplicate passages, changing indexes, incomplete judgments, imbalanced classes, overlapping labels, nested entities, no-answer cases, and high-cost false positives and negatives to compare lexical, rule-based, statistical and learned approaches.
Confirm: The person defines the user task and relevance unit before metrics, keeps held-out queries and documents independent, chooses ranking and classification measures with denominators and error costs, examines coverage and failure slices, distinguishes retrieval from answer quality, calibrates or abstains where useful, preserves cited spans, and can keep a simpler baseline when it is more reliable or operable.
03Grounded generation and evaluation
Ask for a summary, translation, or answer across missing, conflicting, stale, multilingual, restricted, and adversarial sources. Include fluent fabrication, unsupported synthesis, quotation drift, prompt injection, sensitive completion, harmful content, style pressure, a benchmark with clustered items, and human reviewers who disagree.
Confirm: The person separates retrieval, source support, synthesis and final approval; constrains context and tools; preserves source identity and limitations; tests refusal, uncertainty and no-answer behavior; evaluates facts, coverage, relevance, language, style, safety and task utility separately; reports benchmark assumptions and uncertainty; and never treats fluency, one judge, or one aggregate score as proof.
04Production language-system operation
Review input validation, encoding and rendering, document and index freshness, model and prompt versions, retrieval latency, context limits, caching, cost, rate and abuse limits, sensitive logs, multilingual telemetry, feedback selection, harmful and adversarial use, incidents, source correction, rollback, reindexing, retraining, reviewer load, and retirement.
Confirm: The person links outputs to source, index, model, prompt, configuration and release versions, monitors task and language coverage rather than only uptime, protects raw and sensitive text, distinguishes distribution change from proven quality loss, samples field failures with stated blind spots, supports review and recourse, and can pause, correct, reindex, roll back, re-evaluate, replace, and retire the system.