| Internet-Draft | Ad Hoc Semantic Reconciliation | October 2026 |
| Janz & Peters | Expires 9 April 2027 | [Page] |
Interoperation between network-management systems has traditionally required a data model agreed in advance. Where that is insufficient, a richer shared information model or ontology is agreed instead. Both demand broad prior agreement, negotiated and maintained by people. This document examines what becomes possible once the systems on each side can reason. Each side can lift its own data into a complete semantic model. Two such models can then reconcile ad hoc, for the occasion, with no model agreed beforehand.¶
The document develops the semantic model and the two operations upon it: the lift that produces a model, and the reconciliation that bridges two. It gives careful treatment to pragmatics, the layer a schema is least likely to hold and so the part most easily left out of view, yet often one that carries key operative meaning. It considers the form and utility of thin references, separately assessing the two distinct roles they may play. It then reports an empirical study across four network-management scenarios, run over independently constructed cases and several cognitive-agent families. The study measures the portability of a lifted model, the extent to which cognition can reconcile divergent models without support of a pre-agreed standard, where a reference is load-bearing and what it usefully comprises. It closes with conclusions for how information should be prepared for cognitive systems, and for how the community's standardization effort might evolve.¶
This is a research document intended to inform discussion in the IRTF Network Management Research Group. It reflects the authors' ideas, thoughts and experimental findings and does not represent IETF or IRTF consensus.¶
This Internet-Draft is submitted in full conformance with the provisions of BCP 78 and BCP 79.¶
Internet-Drafts are working documents of the Internet Engineering Task Force (IETF). Note that other groups may also distribute working documents as Internet-Drafts. The list of current Internet-Drafts is at https://datatracker.ietf.org/drafts/current/.¶
Internet-Drafts are draft documents valid for a maximum of six months and may be updated, replaced, or obsoleted by other documents at any time. It is inappropriate to use Internet-Drafts as reference material or to cite them other than as "work in progress."¶
This Internet-Draft will expire on 9 April 2027.¶
Copyright (c) 2026 IETF Trust and the persons identified as the document authors. All rights reserved.¶
This document is subject to BCP 78 and the IETF Trust's Legal Provisions Relating to IETF Documents (https://trustee.ietf.org/license-info) in effect on the date of publication of this document. Please review these documents carefully, as they describe your rights and restrictions with respect to this document.¶
Network management is acquiring participants that reason. Autonomous systems will consume management information, act on it, and exchange it with one another. They will bring machine cognition to a plane whose reasoning and design have, until now, lain with human engineers. The information those systems are handed was shaped for particular programs and the junctions between them, not for stand-alone comprehension by cognitive systems.¶
Two software systems that must exchange information rarely fully share a common model of the world. The classical remedy is a data model agreed in advance. Where that falls short, a richer shared information model or ontology is agreed. Both require broad prior agreement on a common model, one that is negotiated and maintained by people and that can be slow and costly to reach. This document takes up the alternative that opens once the systems on each side can reason for themselves. They can reconcile their divergent models ad hoc, for the occasion, machine to machine, with nothing settled beforehand.¶
The central question is therefore not how to write a better shared data model. It is whether any advance agreement is still required or useful once the participants can reason, and, where it may be, in what form. This document argues, and supports with measurement, that a great deal of what was formerly pre-agreed can instead be produced and reconciled on demand by cognition. What remains worth agreeing is thinner, and more broadly useful, than a full shared model.¶
The account rests on one object and two operations, kept carefully distinct throughout:¶
The semantic model is the output of the lift and the input to reconciliation. The two operations are easily bundled together, yet they are distinct. Completing a single model is needed even for one system in isolation, simply as the price of feeding a cognitive consumer. Reconciling two is needed when two systems must meet. This document treats them separately, then relates them.¶
The contributions are four. First, a self-contained account of the semantic model, the lift, and reconciliation for a network-management audience (Section 3, Section 4, Section 6). Second, a focused treatment of pragmatics, a layer that often carries important operative meaning and is least often discussed (Section 5). Third, a clear separation of the two distinct roles of a thin shared reference (Section 7). Fourth, an empirical study across four network-management scenarios that measures portability, reconciliation without a pre-agreed standard, and where a reference is load-bearing (Section 8). The document closes with the conclusions the evidence supports (Section 9) and their implications for standardization (Section 10).¶
This document stands on its own. Three earlier Internet-Drafts by one of the authors explored these ideas at an earlier and less mature stage (Section 11). They are cited as related work and familiarity with them is not assumed here.¶
This is an Informational research document. It defines no protocol and uses no normative keywords. The following terms are used with the specific meanings given here.¶
Everything turns on the model a system's data is lifted into. A data model gives form and some of the denotation. What it defers to specification, to convention, and to the engineer's own knowledge are gaps to fill. The semantic model supplies that remainder. It is semantic: it carries meaning, not merely structure. It is ad hoc: built for the case at hand, with no standard model agreed in advance. And it is complete: it carries, as far as possible, the two layers a bare schema generally does not. Those two layers are pragmatics and provenance, and they lie over and above the ontology and its lexicon. The ontology and lexicon fix what the content is about and how its terms are named. Pragmatics governs how the content is to be used. It is the part a schema is least likely to hold, and, as Section 5 shows, often the one that carries important operative meaning. Provenance records who asserted each fact, by what method, and how firmly.¶
This structure is not invented for the occasion. It follows the long-established study of signs [Morris], which distinguishes three aspects of any representation. Syntax is its form, the rules by which it is constructed. Semantics is what its parts denote. Pragmatics is its interpretative context: how it is used. The correspondence to the vocabulary here is direct. Syntax is form, which the semantic model sets aside, since the same meaning may be written in many forms. Semantics, what the parts denote, is carried by the ontology and its lexicon. And pragmatics keeps its name. Data models are strong on syntax and carry some semantics. What they say very little about, in a form a machine can act on, is pragmatics. Pragmatics is a large part of what a human engineer silently supplies when reading a model.¶
As an aside, the same three parts arrive from a different starting point. Placed on the data-to-wisdom ladder [Ackoff], data are values carried by a data model. Information is data placed in relation and given reference, which fixing an ontology and its lexicon provides. Knowledge is information organized to support inference and action, reached by adding pragmatics. The semantic model, ontology and lexicon plus pragmatics, is therefore knowledge in these terms. Wisdom, the top of the ladder, is not a further document at all. It is cognition itself, the faculty that consumes and generates knowledge. On this reading a data model delivers data, the lift carries data up through information to knowledge, and the cognitive system supplies the wisdom that acts on knowledge. The two framings align.¶
Two operations must be separated at the outset, because they are easily bundled and are in fact distinct operations around the same objects. Completing a representation, producing a semantic model so a machine can comprehend it, is one thing. Reconciling two such models, making them mutually comprehensible, is another. The latter is needed when two systems must meet. Figure 1 shows the relationship. The lift produces one semantic model from a system's data. Reconciliation consumes two, and, where the systems must then exchange native data, yields data-model translators as artifacts of the reconciliation process.¶
+----------+ +---------------------+
| System A |-lift->| Semantic model A |--.
| data mdl | | ontology + lexicon, | \ reconciliation
+----------+ | pragmatics, | \ +-------------+
| provenance | '-> | Correspond- |
+---------------------+ .-> | ence: |
+----------+ +---------------------+ / | aligned |
| System B |-lift->| Semantic model B |--' | meaning + |
| data mdl | | ontology + lexicon, | | translators |
+----------+ | pragmatics, | +-------------+
| provenance |
+---------------------+
¶
Both operations are carried out by cognitive agents, and each is realized as a flow in the ReAct family, in which reasoning is interleaved with acting and observation [react]. The two flows are as distinct as the operations themselves: they differ in what they consume, what they do, what they observe, and what they produce, and the sections that follow set out each in its own terms. A further division is worth drawing. The agent performs some operations in its own cognition, reading meaning from what it already holds; it invokes others because their results are not its to author, so that an observation returns from the system, from a counterpart, or from a virtual experiment. Where each operation falls in that division is set out with the operation itself. The released archive [harness] includes a companion note that walks both flows end to end on worked cases, with the agent's reasoning recorded at each turn.¶
The remainder of the document follows this division. First producing the model (Section 4 and Section 5), then reconciling two of them (Section 6), then the reference that assists both (Section 7).¶
The lift turns a system's data into a complete semantic model. It is the first thing asked of machine cognition, and it is a cognitive act in its own right, not a schema transformation. It has a bar to clear. A lift is adequate when a sufficiently capable cognitive agent can pick the resulting model up and apply it to any end, from the package alone. A model that clears that bar is portable.¶
How much the lifting agent has to work from depends on how it is situated. An agent embedded within the operational system can interrogate the system for configuration and state, consult whatever documentation and standards are available, and run virtual experiments to observe what a concept does, without altering the running system; so situated, it can always produce a portable lift. Lifting from a static data surface alone, with none of that access, is the stringent floor, and the study examines that floor closely.¶
In lifting, the agent holds a model under construction and works toward the adequacy bar above: at each turn it reads the current state of that model, picks the concept or facet least settled, takes one operation, folds the result back in, and repeats until the model is portable or no affordable operation would improve it. The operations it performs in cognition are the surface reading itself, fixing a concept's kind, its relations, and its meaning from the structure, gloss, synonyms, and examples in front of it. The operations it invokes are the ones an embedded agent has and a surface-only agent does not: interrogating the system for configuration and state, consulting documentation and standards, and running a virtual experiment to observe what a concept does. Each invoked operation returns an observation the agent could not have produced by reading alone, which is why access, not effort, is what separates a portable lift from one stalled at the floor. Figure 2 shows the flow.¶
The lift: build one portable model from one system
+-------------------------------------------------------------+
| reason: pick the least-settled concept or facet |<--+
+-------------------------------------------------------------+ |
| |
v |
+-------------------------------------------------------------+ |
| act | |
| PERFORM (in the agent's own cognition): | |
| read a concept's kind, relations, and meaning from the | |
| structure, gloss, synonyms, and examples in front of it | |
| --------------------------------------------------------- | |
| INVOKE (embedded agent only; an observation returns): | |
| interrogate the system; consult docs and standards; | |
| run a virtual experiment on what a concept does | |
+-------------------------------------------------------------+ |
| |
v |
+-------------------------------------------------------------+ |
| observe: fold the result into the model under construction |---+
+-------------------------------------------------------------+ repeat
|
v (portable, or no affordable step remains)
a portable semantic model
¶
A semantic model is a set of lifted concepts, each carried under a short header. It helps to see a single such concept in full before measuring what a reader can do with the whole model. The concept below is drawn from the configuration scenario, an ONF Transport API connectivity-service, and it shows the parts a lift makes explicit:¶
{ "id": "t.cs",
"label": "connectivity-service",
"synonyms": ["service"],
"kind": "service",
"gloss": "an end-to-end connectivity service across the network",
"example": "the A1-A3 ODU2 service",
"relations": [ {"rel": "uses", "target": "t.sip"},
{"rel": "realized-by", "target": "t.cep"},
{"rel": "over", "target": "t.topo"} ],
"instances": ["cs-a1a3-odu2 (ODU2, A1 to A3)"],
"ref": "connection-service",
"pragmatics": {
"use": "the billable, SLA-bearing unit of connectivity",
"authority": "owned by the service layer; the transport side
may not redefine its grade" },
"provenance": {
"asserted-by": "the ONF TAPI controller",
"method": "context export",
"firmness": "authoritative" } }
Read top to bottom, the concept separates cleanly into the layers of Section 3. The label and synonyms are the lexicon. The kind, relations, and instances are the ontology, and they are exactly the typed schema (the TBox) and the populated data (the ABox) that make up the knowledge graph a lifted model contains. The gloss, example, and ref are supports for a reader rather than model content: the canonical example illustrates the concept for a person and is not one of the instances that populate the ABox. The last two blocks are the ones bare schema normally do not hold. The pragmatics record what the concept is for and whose judgment governs it, so that a consumer knows the service is the billable unit and that the transport side may not redefine its grade. The provenance records who asserted the concept, by what method, and how firmly.¶
OWL and RDF can carry the full semantic model. The structural layers map natively, as Table 1 shows; pragmatics rides as annotation, and provenance has a dedicated vocabulary, PROV-O [W3C.PROV-O]. Our JSON holds no representational advantage; the two are alternative serializations of the same content.¶
| Lifted-concept field | OWL / RDF |
|---|---|
| concept | owl:Class |
| label | rdfs:label |
| synonyms | skos:altLabel |
| kind | a class annotation |
| gloss / example | skos:definition / skos:example |
| ref | skos:closeMatch to a reference entry |
| relations | typed object properties |
| instances | owl:NamedIndividual, typed to the concept |
What determines whether the pragmatic layer is usable is not the serialization but how the model is consumed. A formal, deductive reasoner over either form can act on classes and individuals, yet it can only treat a pragmatic rule such as "recomputed from the current context, lower during a maintenance window" as inert annotation, since deduction cannot resolve a defeasible, context-relative judgment. A cognitive reader of either form resolves that same rule, reading it against the live operating context. The two serializations are interchangeable here; the mode of consumption is not.¶
The deeper point is that pragmatics is the content's context of use. Handling it is therefore nothing other than resolving context of use, which only a reader that reasons about context can do. What makes the pragmatic layer tractable is the kind of reader, not the serialization: a model consumed by cognition rather than by deduction. The semantic model is thus expressible in the community's existing knowledge-graph standards [W3C.OWL2] [W3C.RDF11] [W3C.SKOS], and what the approach adds is the reader that resolves its pragmatics.¶
Portability is worth isolating because it can be assessed directly. The test hands a lifted model to an independent consumer agent, one that never saw the source system and holds no inside knowledge of it. The agent explains each concept in its own words, from the package (i.e., from the lifted semantic model) alone, and may answer that a concept cannot be determined from what it was given, since an honest abstention is a valid response and not a forced guess. A separate, fixed judge then grades each explanation by meaning rather than wording, against an authored answer key, marking it faithful, partial, an honest abstention, or invented (a meaning the key does not support). This fraction is the comprehension score (see Terminology): the share judged faithful. The whole assessment is run with a strong, a middling, and a weak consumer from the OpenAI GPT family, so the result is read across a capability ladder rather than at a single point. Figure 4 reports the result on the hardest material used in the study: concepts drawn from two unrelated domains, with no shared vocabulary to lean on.¶
comprehension (meaning score) on the hardest material
strong |##################################| 1.00
middle |############################# | 0.87
weak |######################## | 0.73
+----------------------------------+
0.0 1.00
¶
Three findings hold, and the numbers are worth stating. First, portability is real. A well-lifted model is understood cold. On routine material comprehension sits at the ceiling for every agent, strong, middling, and weak alike (a comprehension score of 1.00 across the capability ladder). Second, it degrades gracefully rather than abruptly as the reading agent's capability recedes. On the hardest cross-domain material the score is 1.00 for the strong reader, 0.87 for the middling one, and 0.73 for the weak one: a soft gradient, not a cliff. Portability is therefore nearly, though not entirely, consuming agent capability-independent. Third, it is a property of the agentic model capability, not the agentic model lineage. A capable reader from a different family (DeepSeek) reads a lift about as well as one from the family that produced it (a comprehension score of 0.97, against 0.99 for the OpenAI GPT strong tier), and it beats the OpenAI GPT weak tier. The comprehension judge does not favor its own family: freezing the answers and grading them with two independent judges returns the same score to three figures (0.900 against 0.900).¶
A caveat is that a sufficiently weak agent can confabulate, reporting an understanding it does not have. A weak open model (Qwen) scored 0.56 cold, and its per-trial results were bimodal, either nailing a case or missing it entirely. Portability is thus a claim about capable readers, and, as much, about the attributes of a soundly lifted semantic model. The two go together: the property lives in the model, and a capable reader is what reads it out.¶
Portability is also a property of the lift. What the lift works from is the source surface: the raw material a system exposes, its labels, its kinds and structural relations, and its instance data, before any meaning is made explicit. A meaning-poor surface is one whose labels and structure carry little recoverable sense, for instance opaque identifiers, or names stripped of the definitions that would say what they denote. Such a surface can impair or break a lift to portability. Holding the lifting agent fixed and lifting from progressively poorer surfaces, comprehension falls from a solid-lift baseline of 0.94 to between 0.64 and 0.74 when the meaning-bearing surface is stripped away. One finding here is sharp. A surface that keeps a concept's name but strips its meaning is no safer than an anonymous one, because a bare name invites the lifter to confabulate a sense for it. A reference consulted during the lift repairs most of that loss, restoring the score to between 0.79 and 0.85. This is the first of the reference's two roles. It supplies denotation the source data leaves too thin to lift well. The role is developed, alongside the second, in Section 7.¶
Two variables underlie the lift and locate it against the alternatives. The first is access: whether the cognition is situated effectively "within" the system it lifts, able to consult the running system, its instances and live state; or un-situated, only seeing the visible surface. The second is product: whether the cognition keeps its understanding internal, for its own use, or externalises a self-standing model for others to consume. The lift is the situated and externalising case. Figure 5 places it against the three others.¶
product
^
| up: externalises a model; down: stays internal
| +---------------------------+---------------------------+
| | Re-documenting from text | THE LIFT |
| | un-situated, externalised | situated, externalised |
| | tidier model, still | writes a self-standing |
| | capped by the text | model |
| +---------------------------+---------------------------+
| | Un-lifted consumption | Operating own system |
| | un-situated, internal | situated, internal |
| | reads the artifact; | understands in context; |
| | the un-lifted baseline | no portable artifact |
| +---------------------------+---------------------------+
+----------------------------------------------------------->
access: un-situated (artifact only) --> situated (consults system)
¶
The contrast that matters is the diagonal. Un-lifted consumption, the un-situated reader that has only the artifact, measures that artifact's portability without the lift; run over a model that was never lifted, it is the un-lifted baseline. The lift is the opposite corner, where a situated cognition exercises various options to complete a portable semantic model. What it adds along the diagonal is access, the reach into the system, and grounded generation, the writing-down of what it finds. The remaining corners isolate each half: a situated cognition that does not externalise is an agent operating its own system, understanding in context but leaving nothing portable; an un-situated cognition that does externalise is re-documentation from the text, which tidies but stays capped by what the text holds.¶
Read against the lift's flow (Figure 2), the access axis is just which of the flow's operations are open to the agent. Situated, on the right, it has the full action set: the invoked operations, interrogating the system, consulting documentation and standards, running a virtual experiment, are all available, and the loop can close gaps the surface leaves. Un-situated, on the left, those system-facing operations fall away; what remains is the performed core, the surface reading itself, so the flow collapses toward a surface-only lift. The access the diagonal adds is, in these terms, the restoration of the invoked half of the loop.¶
This sharpens the limit of the lift. The place a static artifact defeats an un-situated reader, the fact its text never records, is most often not a true frontier but a gap the lift would close, since the fact is present in the system and a situated lift reaches it. What a concept needs then divides in three: what is implicit but present, made explicit by grounded generation; what is present but elsewhere in the system, supplied by situated access; and what is neither, the residue of authority and underdetermination, referred to a person.¶
The quadrangle and the cognition spectrum (Section 6.4) are complementary. The spectrum grades one of the quadrangle's axes, access: where the quadrangle asks whether a cognition is situated or holds only the artifact, the spectrum makes that a dial and splits it per reconciling party, from both live through one inert to both inert. It has no axis for the quadrangle's other dimension, product, which is whether the act externalises a self-standing model or keeps its understanding internal, and that is the distinction that separates the lift from mere comprehension. So the two do not collapse into one: the quadrangle names which act is being performed, and the spectrum grades how far that act reaches as cognition recedes from one or both of a pair of reconciling systems. When access is applied to reconciliation, the access that matters is to the other system, which is what the spectrum's placements grade.¶
Of the parts of a semantic model, pragmatics is the one least talked about. That is probably because it does not freeze cleanly into a schema the way the others do. As it seldom features in schema, it tends to fall out of view and out of consideration. It is, nonetheless, important in operations, and worth setting out plainly. Completing information for a cognitive consumer is, in large measure, capturing the pragmatics a data model leaves out.¶
Where semantics fixes what a term denotes, pragmatics fixes how the denoted thing is to be used. A semantic model records, as first-class content and over and above the ontology and lexicon, matters such as these:¶
Two things about pragmatics emerged clearly from the study. The first is that pragmatics must itself be lifted. It does not as a rule feature in data models, and it is precisely the part a human engineer supplies from outside the artifact. A lift that captures ontology and lexicon but omits pragmatics produces a model that looks complete and is not, and the omission surfaces only when the model must be used to support making operational judgments. The second is that pragmatics is often a highly load-bearing layer. In a range of important operational tasks, signal meaning may be carried as much by pragmatics as by ontology or data structure.¶
Observability is the sharpest case, in the sense of demonstrating both the importance of pragmatics and how it must be handled. The same rising measurement is a matter to act on during normal operation, benign inside a maintenance window, and a matter merely to watch when the detector's confidence is low (Figure 6). None of that is in the value or its type. It is entirely in the pragmatics. The dangerous confusion in this domain is itself a pragmatic one. An alarm, a declared undesirable state, is not an anomaly, a deviation that may mean nothing. Treating them as the same term is a category error that a purely semantic reading does not catch [RFC9940]. The study bears this out with numbers. When the pragmatic content was carried, the strong and middling agents returned the correct significance verdict (accuracy 1.00 and 0.83) and the false-page storm disappeared. When it was withheld, the same agents paged nearly everything (accuracy 0.17 and 0.08). The sound operational verdict could not be judged from structure alone.¶
signal context verdict
------ ------- -------
,-- normal operation -----> act now (page)
a rising |
measurement -+-- maintenance window ---> benign (suppress)
|
`-- detector confidence --> watch
low
¶
Intent shows the same layer deciding a different kind of question. A declarative expectation, such as a lower bound on capacity, an upper bound on latency, or an availability target, is satisfied by a concrete realization rather than equal to it. Which realisations satisfy it, and how conflicts among expectations are traded off, is pragmatic content, not structural. In the study the pragmatic judgment carried the negotiation. A provider returned a best-achievable offer, and the consumer's own declared priorities accepted a degraded value, with no human in the loop (decision accuracy 1.00 for the strong and middling agents).¶
One further finding sharpens where the payoff sits: how hard a pragmatic judgment is to apply depends on whether it must be made live or has been made in advance. Resolved live, from the operating context, it is an act of judgment, and judgment is capability-gated. Handed the significance context, the weak agent barely moved (accuracy 0.53 with the content, 0.50 without), while the stronger agents reached the right verdict (the 1.00 and 0.83 above). Carried instead as a pre-placed movable policy, a context-to-verdict table the live agent applies by lookup (Figure 7), the judgment has already been made, once, and the live reader need not make it again; the capability gate lifts. Both forms are pragmatics: the policy is the significance judgment pre-decided and written down as structure, while the use and authority of the concept in Figure 3 are the same layer left as prose, to be judged against the live context. The cost of the pre-placed form is not avoided but moved: the judgment is paid for once, in authoring the policy, rather than re-incurred by every reader. So a pragmatic judgment made live is gated by the reader's strength; the same judgment pre-placed as a policy is not.¶
Pragmatics as a pre-placed policy: a significance judgment, made once,
applied by lookup rather than judged live
incoming: a rising measurement + its operating context
|
v
the pre-placed policy:
+-------------------------------+-------------+
| operating context | verdict |
+-------------------------------+-------------+
| normal operation, high concern| act (page) |
| normal operation, low concern | watch |
| maintenance window | suppress |
| detector confidence low | watch |
+-------------------------------+-------------+
|
v
verdict read off by lookup: no live judgment, any agent applies it
¶
Because pragmatics is content, it also diverges between two models and must be reconciled like any other layer. Indeed it is one of the principal axes of divergence in Section 6. The practical consequence is a design instruction. A lift that omits pragmatics is not merely incomplete but silently so, and a layer that needs deliberate capture is one a data model is least likely to hold.¶
A semantic model is an object, and reconciliation is a flagship use of it. Reconciliation aligns two independently built models, so that the systems that use them natively can interoperate. It is a distinct operation with its own objective, bridging the genuine divergence between two adequate models. That is not a matter of patching bad lifts. Two sound models still diverge, and bridging that divergence is the task.¶
To see the shape of that task, recall what a lifted model contains. It contains a knowledge graph: a schema of typed concepts and permitted relations (the TBox), and the individuals that populate them (the ABox). Over that graph sit the pragmatics and the provenance. Reconciliation must bridge the divergence between two such models wherever it arises: in the ontology, among the instances, and in the pragmatics. These differ in kind, and a treatment that runs them together will miss half the work.¶
Like the lift (Figure 2), reconciliation is also realized as a ReAct-style flow, but a distinct one. Its input is two finished models, not one system's surface, and its product is a set of confirmed correspondences, with aligned meaning and the translators that follow, alongside an honest record of what stays unmatched. Each side runs the loop, and the two are in contact: at each turn an agent reads the state of the alignment so far, chooses the correspondence most in need of settling, takes one operation, folds the result in, and continues until it judges itself at a confident close or its budget is spent. The operations it performs in cognition are the judgments the subsections below set out: proposing a correspondence, refusing a false cognate, deriving a missing concept, and constructing an entry in the shared reference. The operations it invokes are those whose results are the other party's or the system's to give: eliciting the private semantics a bare catalogue does not carry, and running the decisive experiment that separates genuine look-alikes or confirms a refinement. Those observations come back from the counterpart and from the live system. Read this way, the subsections that follow are the action set of this flow: the kinds of divergence it must settle (Section 6.1), the confirmation step (Section 6.2), and the construction of shared ground where none exists (Section 6.3). Figure 8 shows the flow.¶
Reconciliation: bridge two models (each agent runs this, in contact)
+-------------------------------------------------------------+
| reason: pick the correspondence most in need of settling |<--+
+-------------------------------------------------------------+ |
| |
v |
+-------------------------------------------------------------+ |
| act | |
| PERFORM (in the agent's own cognition): | |
| propose a correspondence; refuse a false cognate; | |
| derive a missing concept; build a shared-reference entry | |
| --------------------------------------------------------- | |
| INVOKE (the result is the counterpart's or the system's): | |
| elicit the private semantics a bare catalogue omits; | |
| run the decisive virtual experiment that separates | |
| look-alikes or confirms a refinement | |
+-------------------------------------------------------------+ |
| |
v |
+-------------------------------------------------------------+ |
| observe: fold the reply or result into the alignment |---+
+-------------------------------------------------------------+ repeat
|
v (confident close, or budget spent)
confirmed correspondences + translators, with residuals
¶
Ontological correspondence establishes how the two schemas correspond, concept to concept and relation to relation. Two models diverge here on lexicon (the same concept under different names, or different concepts under one name) and on structure and span (what each treats as an entity, and how finely). One side may bundle into a single concept what the other separates into several. A single legacy alarm, for example, corresponds one-to-many to a separated alarm-state and the fault it implies. Ontological correspondence is a type-level judgment, and where the relation is a refinement rather than an identity it must be confirmed against the specific relation that holds: a bound is satisfied, not equalled; an underlay is ridden, not renamed.¶
A third kind of divergence is span itself: a concept one model carries and the other simply lacks, with no counterpart at any cardinality. This one-to-zero case differs from a bundled or refined correspondence, where the content is present on both sides and only carved differently; here it is present on one side and absent on the other, because the two models cut the world at different joints. One side reasons in services while the other holds only the resources a service is realized over. Whether the gap matters depends on the exchange. Where the absent concept is not needed, the honest outcome is to leave it uncovered: a capable agent returns it unmatched rather than inventing a partner. Where it is needed, two disciplined routes supply it. The first is to derive it, since the gap is usually a seam between related things: the deficient side can often synthesize the missing concept from what it does hold, as a service is recovered as a binding over the resources that realize it. The second, where derivation is unavailable, is to extend the shared ground rather than either native model, adding one entry to the ad hoc reference both sides bind through; that keeps the addition disposable, whereas writing the concept into a system's own schema is the heavier re-documentation this document sets aside. The constructed-reference case of Section 6.3, an IP service meeting a transport circuit, is exactly such a seam and shows the second route in use.¶
Instance matching, also called entity resolution or co-reference, aligns the individuals once the concepts correspond: which live circuit on one side is which on the other. This is a knowledge-graph-level judgment over keys, attributes, and topology. Most pairs are settled from the models alone. Some are genuine look-alikes, structurally identical on paper, that no static description can separate. Those are settled only by acting, but virtually: a decisive experiment operated on the models rather than the live network, which Section 6.2 takes up.¶
Pragmatics diverges too, and it reduces to neither of these. It is not the ABox. An instance's current state is observed; whether that state warrants action here and now is judged. So pragmatic divergence is settled not by matching but by resolution: by live elicitation of a party's judgment, or by a pre-placed movable policy that carries that judgment to where it is needed. As cognition recedes, pragmatics is the hardest layer to recover, precisely because it is resolved against a live context rather than stored in the graph.¶
The first thing capable cognition does, with no shared reference at all, is bind what clearly corresponds and refuse what merely looks alike. The discipline is the point. A planted false cognate is refused, not bound. Consider a transport grade that is a protection class, set against an IP grade that is a class of service: the same word, unrelated meanings. A capable agent declines the pairing on non-lexical evidence, the concepts' kinds, attachments, and instances, which a matcher keying on the shared label cannot do. A capable agent left without a reference may commit little, but what it commits is correct. It defers rather than guesses. This is the honest behavior, and it is what makes the resolved fraction and precision separate measures, both worth reporting.¶
Reconciliation does not stop at proposing correspondences. It confirms them, as a step of its own, and keeps only what it can confirm. There are three ways to confirm, and the study exercised all three. These are the verification operations of the reconciliation flow (Figure 8): the tools the loop reaches for to confirm a proposed correspondence before keeping it. How much of this verification is available depends on where the cognition sits. The three ways below describe it at full reach, with both reconciling systems live and able to reason. As cognition recedes from one side or both, less of it can be brought to bear, which Section 6.4 develops as the cognition spectrum.¶
The first is question and answer between live agents. Each side is the authority on its own model, so whatever is ambiguous the other can ask about and get an authoritative answer. In one recorded run, two agents reconciling a cross-domain seam asked and answered one question per concept before either proposed a single binding.¶
The second is the decisive virtual experiment. Because reconciliation operates on the models rather than the live network, an agent can provision a candidate correspondence in virtual space, operate it, and read back whether the invariants a correct translation must preserve still hold. The question is then settled by operation, not by argument. In a recorded configuration run, one candidate was provisioned and confirmed as an identity, while another was provisioned and refuted, both sides then marking those concepts as having no counterpart.¶
The third is round-trip and invariant checking for an identity correspondence, and a satisfaction check for a refinement, where a value cannot be round-tripped because the relation is lossy. Verification is what drives precision toward 1.00 across the spectrum: it rejects the false cognates that slipped into a proposal, and it defers the correct-but-unconfirmable pairings rather than asserting them.¶
Where the two models share no reference and none is supplied, the agents can construct a thin shared reference themselves and bind through it. This is the hardest case: two private systems with no public standard between them. Concretely, one such case set a home-grown transport system against a home-grown IP/VPN system, whose worlds touch at a single seam where an IP service rides a transport circuit as its underlay. Making one order across that seam requires binding the circuit to the underlay, binding the transport hand-off to the IP attachment, and pinning three requirements so they cannot be misread: a committed payload rate that is not the bearer line rate, a latency bound that is not a measured value, and protection against a path failure rather than an IP scheduling class. It also requires refusing one look-alike, the transport and IP grade.¶
A reference for this case is deliberately small. It carries one entry per shared seam category, each fixed by a canonical example both sides can instantiate. It does not cover either side's native concepts, and that omission is deliberate: with no grade entry to bind to, the cross-domain grade false cognate is pre-empted at the reference itself. The listing below shows the container and two of its five entries.¶
{
"id": "ref.crossdomain.adhoc.v1",
"kind": "lexical",
"reference_type": "ad_hoc_by_example",
"note": "Constructed between these two agents for this exchange,
not adopted from any standard. One entry per shared seam
category, each fixed by a canonical example. It does NOT cover
either side's native concepts (grade, bearer, vlan, service);
that omission pre-empts the cross-domain 'grade' false cognate.",
"entries": [
{ "id": "transport-underlay",
"label": "transport underlay",
"synonyms": ["underlay circuit", "carrier circuit"],
"class": "seam",
"definition": "The transport circuit that carries an IP
service as its underlay: the one object both domains
share, owned by transport and consumed by IP.",
"example": "circuit CIR-7 is the underlay of service SVC-42" },
{ "id": "service-rate",
"label": "committed service rate",
"synonyms": ["committed rate", "CIR"],
"class": "requirement",
"definition": "The committed client payload rate the service
guarantees, NOT the bearer line rate.",
"example": "5 Gbit/s committed (not the bearer line rate)" }
// + 3 more: demarcation, latency-bound, protection-requirement
]
}
Each entry is thin: a stable identifier, a preferred label and synonyms, a shallow class, a disambiguating definition, and one canonical example. It carries no relationship axioms, no cardinalities, and no attribute schemas. It is a lexicon, not a model of the domain. Figure 10 reports what such a self-built reference buys for the strong agent on this case. Left alone with no shared reference, the resolved fraction is 0.20: the agent refuses to guess the seam and so commits little, at full precision. Constructing the thin reference and binding through it lifts the resolved fraction to 0.80. A single decisive virtual experiment on the remaining candidates closes most of the rest, reaching 0.90, against a reference-given ceiling of 1.00. Precision stays at 1.00 throughout, and no false cognate is taken. The climb is in what the agents surface, not a trade against correctness.¶
resolved fraction (precision stays 1.00 throughout)
no shared reference |###### | 0.20
agents build one |######################## | 0.80
+ one virtual verify |########################### | 0.90
+------------------------------+
0.0 1.00
¶
A resolved fraction below 1.00 here is deferral, not failure. It is what a thin, self-built reference has not yet surfaced, and what the reconciliation's own verification then closes. The reading that matters is that the agents need no standard handed to them. Where one is absent, capable agents can build the ground the reconciliation stands on. The full role of the reference, and its economics, are taken up in Section 7.¶
How far a reconciliation can proceed on its own depends on where the machine cognition sits relative to the two systems. This is the master variable, and it governs what a reconciliation achieves, what it costs, and what it can verify. It runs along a spectrum (Figure 11): both sides cognitive, one side inert, or cognition applied only in overlay to two inert systems. In the terms of the reconciliation flow (Figure 8), the spectrum is just how many of the two loops are live: a full loop on each side, a full loop on one side only, or neither, with cognition then reduced to an overlay that can propose correspondences but cannot run the loop's verification operations in place.¶
cognition recedes ------------------------------>
+------------------+ +------------------+ +------------------+
| Both cognitive | | One side inert | | Cognition in |
| | | | | overlay |
| verify: mutual | | verify: live | | verify: propose |
| Q&A + joint | | agent runs solo | | candidates for |
| virtual | | virtual experi- | | an external check|
| experiments | | ments; reference | | |
| | | supplies facts | | |
| reconciliation | | | | no full autonomy;|
| ALWAYS completes | | mostly completes;| | external |
| (autonomous, | | more is deferred | | adjudication |
| verified) | | | | |
+------------------+ +------------------+ +------------------+
¶
At the fully cognitive end, both sides can reason, so both loops run with their full action set, and here a strong claim holds for sufficiently capable agents. Reconciliation always completes. This is an in-principle fact, and in the study it was also verified. It holds for two reasons, both of them a result of having live cognition on each side. First, the agents can exchange unbounded further information: each is the live authority on its own model, so whatever is ambiguous the other can ask about and get an authoritative answer. Second, the agents can run decisive virtual experiments, since reconciliation operates on the models rather than the live network. Any question is therefore confirmed, refuted, or authoritatively decided. Across the four scenarios the study found no semantic gap that two sufficiently capable agents could not close, including on the two operations that look least automatable: a two-sided intent negotiation, and the significance verdict in observability. What remains after everything has been exchanged and every experiment run is never an unbridgeable correspondence. It is a concept with no counterpart (correctly returned as unmatched), a fact not yet realized in the running network (an absence in the world, not in meaning), or one of the two irreducible residues discussed in Section 9. None of these is a gap cognition gets stuck on.¶
As cognition recedes from one side, reconciliation degrades gracefully rather than breaking. Only one loop is live now: the live agent can probe the static description and run solo virtual experiments but cannot interrogate a peer or co-design an experiment, so its loop keeps the system-facing operations and loses the counterpart-facing ones. It reconstructs the inert side's meaning from structure and instances, and a thin reference buys back the facts the inert side cannot supply. Pragmatics is the hardest layer to recover here, because it was resolved against a live context that is now gone. The effect is measurable and its direction depends on capability. For a strong agent a reference is an effort substitute: given the anchor, its hidden deliberation on one schema-binding task collapsed by a factor of about 3.7 with both sides live, about 6 with one side inert, and about 2.4 with both inert. For a weak agent facing an inert side the same reference can add effort rather than save it.¶
At the far end, cognition sits only in overlay to two inert systems, with neither loop live in place. It can reconstruct both from structure and data, but it can only propose correspondences for external adjudication. Full autonomy is not reached. That boundary is real, and it is stated as such in Section 9. The through-line across the spectrum is one claim, made concrete on measured curves: it is cognition that completes a reconciliation, and where cognition is present in full, the reconciliation completes.¶
A shared reference recurs throughout the account, and it earns its place in two separate ways that are easily conflated. Stating them apart is the point of this section.¶
In its first role, a reference supports the lift. Consulted while a model is being produced, it supplies denotation the source surface is too thin to carry, so that a meaning-poor data model can still be lifted into a comprehensible one. This is the repair measured in Section 4, where a reference used in the lift pulled comprehension on a stripped surface back from the 0.64 to 0.74 range up to the 0.79 to 0.85 range. This role serves comprehension generally.¶
In its second role, a reference supports reconciliation. It is the shared common ground two models bind through. Because each side binds its terms to the reference, genuine correspondences arrive as near-identity rather than as candidates to be argued, and false cognates are pre-empted, since two terms that merely share a label bind to different entries. The residual work is confined to what the reference does not cover. This role serves an operation between two parties.¶
These are genuinely separate roles. A careful lift can make a model comprehensible with no reference at all, so the first role may fall away entirely, while the operations that later use the model can still benefit from the second. The practical conclusion is therefore to make the reference, or a pointer to it, part of the model package regardless. It is thin, and may prove useful to downstream operations.¶
Its needed nature follows from both roles and is deliberately minimal, as Figure 9 showed. We find a sound reference to comprise a lexicon with thin supports. For each entry it carries a stable identifier, a preferred name and synonyms, a shallow class, a disambiguating definition, and a single canonical example. That example is explanatory, an aid to a reader, not an instance; the reference holds no ABox of its own. It carries no relationship axioms, cardinalities, or attribute schemas. It is not a model of the domain, and it is emphatically not a universal ontology all parties must accept. In practice it is most often derived from an existing model or standard, and, as Section 6.3 showed, where none exists capable agents can build one. One caution the study makes precise: what binds correctly is the reference's description, never a bare shared pointer. A shared identifier with no description behind it is worse than no reference at all, because an agent may bind by an opaque token - and wrongly.¶
The usefulness of a reference is also graded. It is load-bearing only where there is no standard to lean on and the agent is capable enough to exploit it. Where a public standard already supplies the common ground, a further reference adds little. Where the agent is too weak to reason over the reference, it cannot rescue the outcome. This shows cleanly on the observability ontological cognate: the RFC 9940-anchored reference rescued the middling agent completely, taking it from always conflating the alarm and anomaly concepts to never doing so, while the weak agent conflated the terms with or without it, and the strong agent never needed it. Between those limits the reference does real work, and it does so cheaply: supplying an apt reference cut the cognitive effort a reconciliation spent, measured in reasoning tokens, by a large factor on the harder scenarios, because the agents argued far less.¶
The reference brings a second, independent economy, this one structural. Because each system binds to a shared reference rather than to every other system pairwise, the operational step count associated with reconciliation grows with the number of systems rather than with the number of pairs. N reconciliations to a shared reference replace on the order of N(N-1)/2 pairwise reconciliations. The saving is a matter of counting; what the study checked, out to twelve systems, is that composing the N single-reference bindings reproduces the correct pairwise correspondences, so the linear path is correct and not merely cheaper.¶
The claims above rest on a study run across four network-management scenarios. The scenarios were chosen so that together they span a range of operation classes and the principal ways two models diverge. Each was built, run against a validated gold standard, and reproduced on independently constructed cases. The method is summarised here only as far as is needed to read the results.¶
Each scenario was run on three independently built cases: three different pairs of semantic models to reconcile, deliberately varied in domain, vocabulary, and planted traps, so that a finding resting on one case could be tested against two more. The agents spanned three independent families, the OpenAI GPT family, DeepSeek, and Qwen, so that a result could be checked against the possibility that it was an artifact of one model lineage. The OpenAI GPT family supplied a three-tier capability ladder, gpt-5.6-sol (strong), gpt-5-mini (middling), and gpt-5-nano (weak); DeepSeek's deepseek-chat and Qwen's qwen2.5:7b supplied a strong and a weak reader from other families. The prose below refers to these by family and tier. Outcomes were scored against a validated gold standard for each case. The two headline measures are the resolved fraction and precision defined in Section 2. Cognitive effort was measured in reasoning tokens.¶
Each scenario's signature result reproduced on its independently built cases and held across the agent families and capability tiers. Table 2 summarises this, scenario by scenario.¶
| Scenario | Signature result | Reproduces across |
|---|---|---|
| Configuration | Capable agents reconcile two standard models on their own (precision and resolved fraction 1.00 for the strong agent); a thin reference mainly prevents the weaker agents' errors. | 3 cases; families; tiers |
| Intent | Refinement and two-sided negotiation complete autonomously at the cognitive end (decision accuracy 1.00); accuracy over the full multi-step lifecycle falls with agent capability (1.00, 0.88, 0.62). | 3 cases; families; tiers |
| Standard-free | Without a shared reference the strong agent under-commits at full precision (0.20); a constructed reference completes the close (0.80, then 0.90 with a decisive experiment). | 3 cases; families; tiers |
| Observability | The alarm-versus-anomaly false cognate is refused, and the significance verdict is correct once pragmatics is carried: verdict accuracy for the strong and middling agents is 1.00 and 0.83 with it, 0.17 and 0.08 without. | 3 cases; families; tiers |
Two invariances are worth stating on their own. Every headline result reproduced across the independently built cases, different model pairs returning the same finding, so the results are not a property of one hand-built example. And the headline results held whichever agent family did the reconciling, so they are not an artifact of one vendor's reasoner.¶
One caution belongs beside the headline numbers, because it bears on acting on a result unsupervised. The dangerous failure is a confident wrong answer, and it concentrates at the weak end and the inert placements. The rate of confident errors rose about elevenfold down the capability ladder, from 0.004 for the strong agent to 0.046 for the weak one, and calibration worsened in step. An agent's own confidence is a weak guard: confidence on wrong bindings sat in the 0.73 to 0.85 range, only a little below confidence on correct ones. A downstream system cannot separate right from wrong by the confidence number alone. The discipline a weak or inert-facing reconciliation needs is an explicit abstain path, a scored declaration of insufficient evidence reported apart from honest deferral, rather than a confidence cutoff.¶
The benchmark is released as a citable, openly available archive [harness]: the four scenarios and their independently built cases as plain-JSON model pairs, each with a key a script derives and consistency-checks before any run; the reconciliation harness; and the non-cognitive descriptor baseline. The released runs are instrumented: each agent records, at every turn, its reading of the state, the correspondence it most needs to settle next, the decisive experiment it will run, how close it judges itself to a confident close, and whether to continue, stop, or escalate. These make the cognitive process this document describes inspectable turn by turn, not only in its outcomes, and a check confirms the record is observational, leaving the agent's decisions unchanged. The agents run as a ReAct-style loop, reasoning interleaved with tool-acting and observation [react]; a companion note in the archive reads these traces turn by turn.¶
Three conclusions about the concept are supported by the evidence.¶
Semantic-model portability is real, and it is measurably verifiable. A capable cognitive agent understands and uses a soundly produced semantic model from the package alone. This can be checked directly, across a capability ladder and across independent agent families, rather than inferred from a downstream task. Portability is a property of the model, and, as much, of the reader's cognitive capability.¶
Capable cognition reconciles divergent semantic models on its own, with discipline. Where the systems reason well enough, two divergent models are reconciled through the agents' own questioning and virtual verification, with no standard supplied. True correspondences are committed and false cognates refused. At the fully cognitive end of the spectrum reconciliation always completes: an in-principle assertion verified by the study.¶
Portability and reconciliation are graded, not all-or-nothing. Both track the cognition available, degrading gracefully as it recedes rather than collapsing. A thin reference can then partly compensate, rescuing a cognitively weak consumer and making a reconciliation cheaper, though only within the bounds Section 7 describes.¶
The limits are equally real and are stated plainly. Two residues do not yield to cognition, because they are not questions of fact. One is authority. When two sides genuinely conflict over whose value governs a contested field, a reference can supply information but never authority, and the decision stays with a human or an owning party. The study drew this line sharply in the intent scenario: publishing an invariant floor lifted a blind agent's satisfaction check from 0.29 to 0.71, but at both sides inert no reference could move a negotiation whose missing ingredient was the customer's own judgment. The other residue is genuine underdetermination. Where no fact settles a correspondence, an honest agent defers rather than invents one. A reconciliation that reaches these residues and defers is behaving correctly, not failing.¶
And there is the autonomy boundary of Section 6.4. Full, unsupervised reconciliation is available at the fully cognitive end. It is not available where machine cognition is employed only in overlay to two inert systems, which can only propose correspondences for external adjudication. Between those ends the automation degrades predictably, and a thin reference buys back part of what receding cognition gives up.¶
If portable semantic models can be produced and reconciled on demand by cognitive systems, the question for the community is not whether to abandon standard data models but whether such models should be evolved, and in particular, whether it might be helpful to introduce other standardized artifact types. Our findings suggest an evolution rather than a rupture.¶
What may become most useful to agree in advance may narrow and simplify. As models can increasingly be produced and reconciled ad hoc, for the task at hand, the heaviest fully-specified schemas may give way to lighter, more broadly usable artifacts. Thin shared references, carrying names, definitions, and examples derived from existing models, are the anchor to which independently built systems bind. Such references are relatively easy to maintain and evolve. They are also precisely the kind of interoperable anchor to which emerging knowledge-graph and AI-modeling efforts in the community could bind. The community's knowledge-graph work for network operations [I-D.mackey-nmop-kg-for-netops] meets this approach squarely, since a lifted semantic model contains a knowledge graph a machine can consume, and a thin shared reference is exactly the kind of anchor such graphs can bind to.¶
What still needs agreeing is narrower than a full shared model, and it is of two kinds. The first is the thin references themselves: the small shared anchors reconciliation binds through, worth agreeing precisely because they are cheap and broadly reusable. The second is the authority boundaries the study keeps reaching (Section 9): the contested fields where two systems hold conflicting values and no reasoning can decide whose governs, because the question is one of authority, not of fact. This second kind is not a modeling problem, and better cognition will not dissolve it. Agreeing who governs what is a matter of governance, and it is exactly where standardization remains indispensable, even as ad hoc reconciliation takes over much of the rest.¶
A further implication follows from where the model now comes from. A semantic model is produced by a system's own cognition, the lift, and its form is flexible: what matters is not a fixed schema but that the model carries the elements this document describes, its ontology and lexicon, its pragmatics and its provenance, and explains itself well enough for a capable agent to pick it up. The burden that once fell on a shared model therefore shifts onto a system's ability to describe and interrogate itself. Where a system exposes enough of its structure, its local concepts, its current facts, and its authority and provenance, a consumer's cognition can lift and reconcile against it ad hoc. Where a system exposes none of that, and the meaning is neither in its schema nor available elsewhere, autonomous reconciliation cannot be guaranteed, and the missing content must be drawn from another authoritative artifact, elicited from a live interface, or supplied by a person.¶
This points to a second candidate for standardization, alongside the thin references: not a universal domain model, but a system's self-description and interrogation. What is worth agreeing is how a system exports its structural surface, explains its local concepts, makes its current facts discoverable, declares its authority and provenance, and answers open questions from a consumer that does not know its schema in advance. One architectural pattern that realises this is an interface that declares the concepts and relationships a system supports and assembles the requested knowledge at runtime, rather than fixing all exchanged content at design time, carrying its own provenance and confidence and able to defer where the evidence is insufficient. The point is the capability, not any one serialization or interface: agree how a system explains itself, not the model it uses inside.¶
None of this displaces the installed base. Ad hoc production and reconciliation will coexist with legacy standards and legacy systems, machine to machine, for the task at hand, while the standardization effort is redirected toward the anchors and the authority boundaries that actually need agreeing. The direction is an evolution of where the community spends its modeling effort, not a repudiation of the value that pre-agreement has delivered.¶
One important next step is to move the study from purpose-built cases to larger, real, operator-sourced datasets, spanning the same scenarios at production scale and with the messiness of real data. Real models diverge in ways a designed case cannot fully anticipate. The measures reported here, portability, resolved fraction and precision, the reference's graded value, and the cost in reasoning effort, should be re-established on operator data before they are relied upon.¶
A further step is more consequential, and only an operator can take it: to attach cognition to a live operating system and have it perform the lift in place, with real access to state, instances and pragmatics, rather than over a frozen export. This matters more than examining further inert data. An inert dataset, however large, fixes cognition at the un-situated corner of Figure 5: it refines the baseline and re-measures the same surface bound, and it cannot separate a fact the lift could recover from the system from one that is genuinely absent, because past the surface everything looks like one wall. A situated lift makes that separation, and what remains unresolved after a maximally situated lift is the true residue, the authority and underdetermination remainder that bounds automation and that this work defines but does not yet measure.¶
What the experiment involves can be stated concretely. The target is a live management plane with a machine-readable operational view: an ONF TAPI controller, an IETF-YANG datastore reached over NETCONF or RESTCONF under NMDA, a gNMI-served state stream, or a controller such as ONOS or OpenDaylight. The cognition is a capable agent given read-only tools over that plane, datastore reads on the running and operational datastores, topology and inventory reads, a context export where the interface offers one, a telemetry subscription, and a bounded budget of read-only queries it may issue to settle a question, together with its prior over modeling conventions and, optionally, a thin reference. Safety is a first-class constraint: read-only credentials, rate limits, an audit of every call, and no path to a write. Situatedness is a dial the experiment should sweep, from the frozen export alone (the un-lifted baseline), through structural and relational reads, to instances read from the live datastore, to live state, context export and bounded probes, and finally to a self-lift in which the system's own agent authors its meaning.¶
For a chosen slice of the model the agent proceeds concept by concept: it enumerates the concepts from the schema surface; reads each concept's structure, relations and real instances; where the source gloss is thin or absent, generates a grounded gloss from structure, instances and prior and checks it against them; reads the pragmatics, what the thing is for and whose authority governs it, from a context export, from governance or configuration, or, where only a human holds it, by asking; records provenance, marking each fact's method and firmness; and proposes a reference binding. A verifier pass strikes any generation the material does not support, and every fact the agent can neither ground nor read is flagged as a residue candidate for an authority to settle. Portability is then judged exactly as in Section 4.2, but on a model lifted in situ: the lifted model, and only the lifted model, is handed to an independent consumer, and the headline number is the value of the lift, the portability of the lifted model minus the portability of the un-lifted baseline on the same concepts. Crossing the situatedness dial with this measure is the core result: as access climbs, portability should rise and the unresolved set should shrink toward the residue, so that what remains at the top of the dial is the first measurement of the irreducible remainder.¶
Two scenario cases carry the experiment. The first is a configuration seam: a standard-governed model, an ONF TAPI or IETF-YANG view, lifted in situ from a live controller. It is the readiest clean case, where public standards are present, the surface is rich, and portability is the headline a situated lift can establish with least friction. The second is the observability case, and it is the sharper of the two, because the pragmatic layer is decisive there. In the study, the significance verdict, whether a rising measurement is an alarm, a benign maintenance effect, or a low-confidence reading, is recovered only when the pragmatics are present (verdict accuracy 1.00 and 0.83) and collapses when they are stripped (0.17 and 0.08), and the context that fixes significance lives in the running system rather than in any static model. A focused run would attach cognition to an operator's alarm-and-anomaly system, lift its alarm model in situ with the live context, bind it to the shared anomaly-semantics reference, and measure both portability and the pragmatic verdict against the un-lifted baseline. The NMOP work on network anomaly detection and its semantics ([RFC9940] [I-D.ietf-nmop-network-anomaly-architecture]) is already assembling the operational data, the term ladder, and, through its hackathons, the venue such a run would need, and the knowledge-graph representation the data can be held in ([I-D.mackey-nmop-kg-for-netops]) is the form a lifted model naturally takes. The barrier to entry is low, because those inputs and that venue already exist.¶
The authors invite operators and NMRG and NMOP participants to contribute real model pairs and datasets, to attempt situated lifts in these settings, and to help define a shared benchmark over them.¶
Beyond scale, several questions the study opens remain. There is lifting from raw schema text rather than a curated surface, and from a cold start with no instances to read. There is the maintenance of a lifted model and of a shared reference over time. And there is the shape of a published reference ecosystem, and how references derived by different parties from the same standard relate.¶
This document describes a direction and reports experimental findings. It defines no protocol and introduces no new wire format. Nonetheless, several considerations follow from the direction. A semantic model and any accompanying reference become inputs a cognitive agent acts on, and so are targets for manipulation. A poisoned reference, or a lifted model with falsified pragmatics or provenance, could steer a reconciliation to an incorrect but confident result. The provenance layer of a semantic model, who asserted a fact, by what method, and how firmly, is therefore security-relevant and should be integrity-protected, and references should be authenticated to their derivation source. The autonomy boundary of Section 9 is also a safety boundary. Decisions of authority, and correspondences an honest agent defers, must not be silently resolved by an agent acting beyond its competence. This concern is sharpened by a measured finding: agents assert wrong bindings at nearly the confidence of correct ones, so a downstream gate cannot rely on a confidence score alone. Reconciliation outcomes intended to drive operational change should carry the confidence and provenance on which they rest, together with an explicit declaration where evidence was insufficient, so that a supervising system or operator can gate action accordingly.¶
This document has no IANA actions.¶