The content types you must prioritise for global release are: software/UI strings, technical documentation and manuals, legal and regulatory documents, e‑learning and training materials, video subtitles and voice assets, product descriptions and commerce copy, customer support and knowledge base content, AI training and annotation datasets, and app/game assets.
Start here before you brief any vendor:
- Software/UI strings — JSON, .resx, iOS .strings files; high update frequency, compliance-critical for regulated apps
- Technical documentation/manuals — DITA XML or Markdown source preferred; large volume, block-segmented
- Legal/regulatory documents — contracts, compliance notices, data-processing agreements; accuracy is non-negotiable
- E‑learning and training — SCORM packages, MP4 video, narration scripts; requires SME sign-off
- Video subtitles and voice assets — SRT/VTT/STL files, raw audio or MXF; timecodes and frame rates matter
- Product descriptions and commerce copy — PIM exports, HTML/XML feeds; high revenue impact, fast cadence
- Customer support/knowledge base — HTML, Zendesk or Confluence exports; directly affects CSAT scores
- AI training and annotation datasets — labelled text, audio transcripts, image tags; quality gates are strict
- App/game assets — strings, UI graphics, in-game dialogue; cultural adaptation often required beyond translation
Three named standards shape how you procure all of it. DITA XML (Darwin Information Typing Architecture) is the structured authoring format that lets documentation be reused and updated without retranslating unchanged content. ISO 17100 is the international standard for translation service providers, covering translator qualifications, review processes, and project management. The UK Information Commissioner’s Office (ICO) enforces UK GDPR, which applies the moment personal data appears in source files you share with a vendor.
What does each content type actually contain?
Effective B2B localisation starts with an audit: you cannot scope a project accurately until you know what you have. Here is what each type looks like in practice.
Software/UI strings are short, context-dependent phrases extracted from code. Typical artefacts: JSON resource files, Android strings.xml, iOS .strings, Windows .resx. Volume is usually measured in strings for software, with mid-size apps typically having a range from dozens to several thousand strings, rather than word counts. The owner is typically the engineering or product team.
Technical documentation is larger, paragraph-segmented, and often authored in DITA or Markdown. A single product manual can be very large and run into many tens of thousands of words across numerous topics. Structured authoring and translation memories prevent expensive rework when content changes post-translation. The content owner is usually technical writing or product marketing.
Legal and regulatory documents include contracts, terms of service, regulatory submissions, and compliance notices. Word counts vary significantly, and accuracy requirements are absolute. These require certified translators and, in many cases, back-translation for validation.
E‑learning and training combines text scripts, SCORM packages, narration audio, and MP4 video. The localisation scope spans text, voice, and sometimes on-screen graphics. SME reviewer involvement is mandatory for regulated sectors such as pharma or financial services.
Product descriptions and commerce copy arrive as PIM exports, XML feeds, or spreadsheets. Cadence is high: seasonal updates, new SKUs, and promotional copy can trigger weekly translation cycles.
Customer support and knowledge base content lives in platforms like Zendesk, Confluence, or Salesforce Knowledge. It is high-volume and fast-moving, which makes TM leverage critical to keeping costs under control.
AI training and annotation datasets are a newer category. They include labelled text corpora, audio transcripts, and image-tagging instructions. Quality gates here are strict because errors propagate into model training.
App/game assets combine UI strings with in-game dialogue, achievement text, and culturally sensitive imagery. Cultural adaptation often goes beyond translation.
Pro Tip: When briefing a vendor, include a sample of 200–500 source segments from each content type. It lets the vendor assess complexity, flag encoding issues, and give you a realistic TM leverage estimate before you commit to a full project.
How do strings and documents behave differently in localisation?
The single biggest technical distinction is string-based vs block-based segmentation, and it changes everything: price, tooling, turnaround, and risk.
UI strings are short, isolated, and context-dependent. A button label like “Submit” has no surrounding sentence to help a translator infer meaning. Externalising strings into resource files with descriptive keys (e.g. button.submit.label) and supplying screenshots gives translators the context they need. Without it, mistranslations are common and expensive to fix post-release.
Documentation segments are longer, carry more context, and benefit heavily from translation memory. When a paragraph is updated, only the changed sentences need retranslation; unchanged segments are reused from the TM at a fraction of the cost.
| Dimension | String workflow (UI/software) | Document workflow (DITA/Markdown) |
|---|---|---|
| Segment type | Short, isolated strings | Paragraphs, sentences, blocks |
| Context source | Screenshots, string IDs, comments | Surrounding text, topic structure |
| TM leverage | Lower (short strings match less) | Higher (repeated paragraphs match well) |
| Primary risk | Missing context, truncation | Layout reflow, outdated segments |
| Typical tools | TMS with string editor, i18n libraries | CAT tool, DITA-aware TMS |
Pro Tip: Run pseudolocalisation before sending strings to translators. It simulates translated text (longer strings, accented characters, RTL direction) and surfaces truncation, overflow, and encoding issues in your UI before a single real translation is produced.
What drives cost and how do you estimate realistic timelines?
Cost is not just word count. These are the real drivers:
- Volume and segment count — words for documents, strings for software
- Update frequency — weekly product copy costs more annually than a one-off manual
- TM leverage — high repetition in documentation can reduce payable words by 30–60%
- File complexity — binary formats (InDesign, SCORM, MXF) require engineering prep time
- Multimedia length — one minute of subtitled video typically requires 100–130 words of subtitle text plus timecode work
- Quality level required — raw machine translation (MT), machine translation post-editing (MTPE), or full human translation each carry different costs and turnaround times
Most teams use a hybrid approach: MT for low-risk support content, MTPE for product UI, and human translation for legal or regulated text.
Timeline is driven by engineering readiness (is the content i18n-ready?), file preparation, SME review cycles, multimedia production windows, and legal sign-off for regulated documents. A rough worked example: a mid-sized UI of several thousand strings with associated resources might deliver within about a week. A large documentation set with accompanying training videos typically requires multiple weeks, depending on complexity and review cycles.
Which file formats and tools actually reduce your costs?
Source format is the single biggest lever you control before briefing a vendor. TMS connectors and pseudolocalisation in CI prevent the hard-to-fix UI errors that inflate post-delivery costs.
Preferred source formats:
- DITA XML — topic-based, reusable, TMS-connectable; the gold standard for technical documentation
- Markdown — lightweight, version-controllable, widely supported by CAT tools
- Externalised JSON/YAML — clean string extraction with descriptive keys for software
- Well-structured Word/Excel — acceptable for legal and support content when DITA is not in use
Workflow components to require in your RFP:
- TMS connector to your CMS, Git repository, or content platform
- Translation memory with match-band reuse reporting delivered per project
- CAT file exchange in XLIFF 2.0 or equivalent open standard
- Automated build and test hooks for software localisation
- Pseudolocalisation test results before pilot sign-off
For website localisation and CMS connectors, confirm the vendor supports your specific platform (WordPress, Contentful, Drupal, Salesforce) before signing. Before you send a single file, run a content readiness audit to flag encoding issues, untranslatable strings, and missing metadata.
Pro Tip: Require vendors to deliver TM leverage reports and pseudolocalisation test results as part of pilot acceptance criteria. If they cannot produce these, they are not set up to reduce your recurring costs.
How do video, subtitles, and voiceover differ from text localisation?
Audiovisual localisation is a different discipline. Get the scoping wrong and you pay twice.
- Subtitles are text overlays timed to speech for hearing audiences; they do not include speaker identification or sound descriptions
- Captions (SDH) include speaker identification and non-speech audio cues; required for accessibility compliance under UK law
- Transcripts are untimed text documents; useful for search indexing but not for on-screen display
- Dubbing replaces the original audio track; highest cost, longest lead time, requires lip-sync adaptation
File and timing requirements: always supply source video with embedded or sidecar timecodes, confirm frame rate (24, 25, or 29.97 fps), and specify delivery format (SRT for web, VTT for HTML5, STL for broadcast, MXF for post-production). Missing timecodes are the most common cause of subtitle project delays.
Common pitfalls to avoid:
- Using raw MT output for customer-facing voiceover scripts (errors are audible and damage brand trust)
- Supplying video without a spotting file or script, forcing the vendor to transcribe before translating
- Ignoring SDH requirements for regulated or public-sector content where accessibility is a legal obligation
- Mismatched frame rates between source and delivery, causing subtitle drift
SME checks on voiceover scripts are non-negotiable in regulated sectors. A mistranslated dosage instruction or legal disclaimer in a training video carries real liability.
What do regulated sectors need from localisation quality and compliance?
Compliance is not a checkbox. It is a contractual and operational requirement that must be built into the vendor relationship from day one.
- ISO 17100 certification confirms the vendor uses qualified translators, a documented review process, and project management controls; request the certificate, not just a claim
- LQA (Localisation Quality Assurance) is a structured error-classification workflow; require a minimum sample rate (typically 10% of deliverable volume) and a written LQA score report per project
- SME reviewer sign-off is mandatory for medical, legal, financial, and defence content; name the reviewer in the SOW and agree their availability before project start
- NDA and secure file transfer — all source files containing personal data trigger UK GDPR obligations; the ICO expects you to have a data-processing agreement (DPA) with any vendor handling personal data on your behalf
- TM change logs and versioned deliverables provide the audit trail regulators and internal compliance teams expect
Pro Tip: Include a liability clause for mistranslation in regulated documents. Specify the error-classification threshold at which the vendor is obligated to rework at no charge, and agree a turnaround SLA for corrections.
| Requirement | What to ask for | Why it matters |
|---|---|---|
| ISO 17100 | Current certificate, not self-declaration | Confirms translator qualifications and review process |
| LQA report | Per-project, minimum 10% sample | Provides auditable quality evidence |
| DPA / GDPR | Signed before file transfer | ICO obligation for personal data in source files |
| TM change log | Delivered with each project | Supports version control and audit trails |
Your practical RFP and vendor checklist
Paste this into your next RFP or vendor interview.
Document-level checklist:
- Source files in preferred format (DITA, Markdown, JSON, XLIFF)
- Glossary and termbase (TBX or Excel)
- Existing TM (TMX format)
- Style guide and tone-of-voice notes
- Screenshots or in-context reference for UI strings
- Character limits per string (for UI)
- LQA criteria and error-classification scheme
- Security classification of source content
Vendor capability questions:
- Which TMS do you use, and which CMS connectors do you support natively?
- Can you provide TM match-band pricing and a reuse report per project?
- Do you have DITA-aware workflows and structured authoring experience?
- What is your ISO 17100 certification status and scope?
- How do you handle SME reviewer coordination and sign-off?
Security and compliance questions:
- How do you handle UK GDPR and ICO obligations for personal data in source files?
- Do you use encrypted file transfer and encryption at rest?
- Are translators subject to background checks for sensitive or classified content?
- What are your data-retention and deletion policies post-project?
Pilot and acceptance criteria:
- Propose a paid pilot of 500–1,000 words or 50–100 strings
- Required pilot deliverables: bilingual review file, TM export (TMX), LQA score report
- Define rejection criteria (error rate threshold) and repair SLA upfront
Pro Tip: Request a pre-project technical audit before the pilot. A good vendor will flag i18n issues in your source files, missing string context, and encoding problems before they become billable rework.
glocco®’s take on prioritisation, tooling and compliance
At glocco®, we approach every new project with the same three-question filter: what is the compliance risk, what is the revenue impact, and how often does this content change? That order matters. A mistranslated regulatory notice in a fintech app is a liability before it is a missed sale. A product description that never updates is a lower priority than a knowledge base that changes weekly.
DITA and TM policies pay back fast, and we have seen this consistently across clients in legal, medical, and enterprise software since glocco® was founded in 2014. The upfront investment in structured authoring and a well-maintained TM typically reduces recurring translation costs materially within the first two or three project cycles. For UK-regulated clients, we layer ISO 17100-aligned workflows, signed DPAs, and LQA reporting on top of that foundation, so compliance is built in rather than bolted on.
Speed, quality, and compliance are not a triangle where you pick two. With the right tooling and the right people, you get all three. That is not a sales line; it is what structured workflows actually deliver.
glocco® is ready to scope your next localisation project
glocco® covers the full range of services this article covers: translation, localisation, interpretation, subtitling, and AI annotation for regulated sectors, across 76 languages and sectors including legal, fintech, pharma, defence, and enterprise software. Every project includes a signed NDA, a data-processing agreement for UK GDPR compliance, and LQA reporting as standard.
Not sure where to start? A paid pilot is the lowest-risk way to validate quality, TM leverage, and workflow fit before committing to a full programme. glocco® also offers a pre-project technical audit to flag i18n issues and file-format problems before they cost you time. For document translation and AI-assisted translation workflows, the team is ready to scope a project with you now.
Request a tailored quote at Glocco.
Sources
- Localize documentation – Globalization | Microsoft Learn
- Building An Effective B2B Content Localisation Strategy | Forrester
- Technical localization: Key steps and best practices
- Software localization and software translation guide
