This study evaluates the translation capabilities of selected AI language models in a defence and security context through a multistage multilingual workflow combining translation, language identification, and back-translation. Using 32 English inputs across eight target languages, the study compares offline and online models in terms of translation quality, processing stability, and task performance across varying levels of sentence complexity. The findings indicate that all models perform strongly on basic translation tasks, but notable differences emerge in language identification, back-translation quality, and output stability. The results help clarify the conditions under which AI translation can support security-relevant multilingual workflows.
This study proposes a multi-criteria framework for automated evaluation of scientific language quality using Large Language Models (LLMs). The approach models translation assessment as a structured decision process across semantic, grammatical, and terminological dimensions aligned with ISO 5060:2024. Validation on biomedical data shows strong agreement with expert judgment (κ = 0.81), while metric-based estimation demonstrates weak correlation. The results indicate that interpretable LLM-based evaluation can reliably replicate expert reasoning and support scalable, transparent language quality assessment.
This paper examines whether adapting a multilingual large language model (LLM) to clearly pro‑Kremlin or pro‑Western texts can improve the automatic detection of pro‑Kremlin narratives in Czech‑language content. The study compares three variants of the same model: a base multilingual model, a pro‑Kremlin adapter trained on Russian‑language texts with pro‑Kremlin framing, and a pro‑Western adapter trained on Western sources that respond to the same topics from an opposing perspective. All models are fine‑tuned using a parameter‑efficient method and evaluated on a small Czech corpus covering five key narrative types. The pro‑Kremlin adapter shows the strongest ability to distinguish between texts with pro‑Kremlin framing and neutral texts, while the pro‑Western adapter brings only a modest improvement over the base model. These findings suggest that exposing an LLM to ideologically aligned training data can make it more sensitive to that ideology, with potential applications for supporting the detection of hostile information operations in smaller‑language NATO member states.
This paper examines how large language models can be embedded in governable agentic workflows for defence decision support. It argues that reliability and accountability in such systems depend on workflow architecture, not only on model capability. Building on the author’s earlier I→E→R model, the paper proposes an audit-first governance framework with three layers: explicit task decomposition, evidence grounding with provenance, and bounded control with logging and authorization gates. The framework is specified through five inspectable roles, four design variables, and bounded autonomy as the guiding design principle. The contribution is conceptual: it defines conditions under which agentic analytical processes can remain transparent, traceable, reconstructable, and auditable, and prepares the framework for later case-study or proof-of-concept evaluation.