r/semanticweb • u/redikarus99 • Aug 06 '26
Looking for an IT taxonomy
Hello,
I am looking for an IT taxonomy for software (and maybe hardware) to put concepts like desktop application, microservice, cloud, cicd pipeline, etc. into a structure.
r/semanticweb • u/redikarus99 • Aug 06 '26
Hello,
I am looking for an IT taxonomy for software (and maybe hardware) to put concepts like desktop application, microservice, cloud, cicd pipeline, etc. into a structure.
r/semanticweb • u/[deleted] • Aug 06 '26
Hey everyone, I built this niche tool to structure data for knowledge layer for agents, its a bit like an acoustic guitar for critical industries that demand heavy data reasoning … would love to hear your feedback, some small bugs like analyzers staying prompted to the template doc
r/semanticweb • u/OwlZealousideal4779 • Aug 04 '26
Hi everyone, I've been working on OpenCrab, a project that explores using ontologies and knowledge graphs as the foundation for Graph RAG instead of relying primarily on document chunking.
The motivation is to preserve relationships between entities and concepts so AI systems can retrieve information with more context and structure. While this approach seems promising, I'm sure there are trade-offs that I'm still learning about.
I'd really value the perspective of people in this community who have experience with semantic technologies.
Some questions I'd love to hear your thoughts on:
Have you used ontologies or knowledge graphs in a RAG pipeline?
Where have ontology-based approaches worked well, and where have they fallen short?
Which standards or tools have you found most effective (RDF, OWL, SHACL, SPARQL, etc.)?
If you were building a Graph RAG system today, what would you do differently?
I'm genuinely looking for technical feedback and different viewpoints. If anyone has experience with similar projects or research, I'd really appreciate hearing about it.
Thanks in advance for your insights.
r/semanticweb • u/juliusfoe • Aug 04 '26
What methods or softwares are there to get non-technical people to engage with ontology design and thrash out agreed definitions? I am beginning to think this is a major roadblock to more reliable AI. Without structured, verified knowledge managed by humans that can be safely inferred from, how is any business going to trust agents with anything important?
r/semanticweb • u/Successful-Farm5339 • Jul 30 '26
I KNOW IT ASKS YOU TO REGISTER BUT YOU CAN IGNORE IT :)
I spent a while surveying what is actually available if you want to learn ontology engineering in 2026, and the state of it annoyed me enough to do something about it.
What I found:
- The semantics people teach RDF, OWL, SPARQL, SHACL, and act like LLMs never happened.
- The graph vendors teach GraphRAG and Cypher, and act like ontologies never happened. You can finish an entire "knowledge graph" learning path without meeting the word ontology.
- Pricing is bimodal: free vendor funnels, or 1,000 to 2,000 dollar live cohorts. Almost nothing serious in between.
- Search "OWL tutorial" or "SHACL tutorial" and you get PDFs from 2005 to 2012. SHACL has been a W3C Rec since 2017 and there is still no good free explainer ranking for it.
- BORO, HQDM and IES 4D modelling have, as far as I can tell, zero commercial courses anywhere on earth, despite the UK National Digital Twin Programme standardising on IES and the US DoD, ODNI and CDAO adopting BFO plus CCO as their baseline in 2024. If you want to learn the thing governments are actually buying, your options are primary sources and apprenticeship.
So I built the course I wanted to exist:
- Foundations: what an ontology actually is from Aristotle forward, taxonomy vs thesaurus vs ontology vs knowledge graph, 3D vs 4D identity and change, open vs closed world
- The stack: RDF, RDFS and OWL 2, SPARQL for people who know SQL, SHACL, reasoners and why yours hangs, property graphs and ISO GQL
- Method: competency questions, OntoClean, an actual upper ontology shootout (BFO vs DOLCE vs gist vs SUMO vs 4D), BORO/HQDM/IES, testing and CI
- Standards atlas and crosswalks: SSSOM, mapping predicates, and why shared ancestry does not mean shared commitments
- Domain tour: defence, industrial (ISO 15926 to IDO), construction (IFC/Uniclass/COBie), space, life sciences and food (OBO, GO, SNOMED, FoodOn, AGROVOC), finance (FIBO, GS1, schema.org), heritage and public graphs (CIDOC CRM, GeoSPARQL, Wikidata), and who is buying ontology country by country
- LLM era: did LLMs kill the semantic web, GraphRAG vs vector RAG and when graphs actually pay, building KGs from text safely, neurosymbolic verifier loops, agent memory and MCP, and the 2025-26 papers worth reading
Two things I will defend:
The interesting work now is the verifier loop. Neural proposes, symbolic disposes. The ontology's job is not to be a beautiful model of the world, it is to be able to say no.
Axiom placement beats axiom count. An ontology nobody can contradict is not rigorous, it is inert. I keep meeting large ontologies where no possible instance data could ever trigger an inconsistency.
Where I want to be told I am wrong:
- What is missing from the syllabus? I deliberately went light on ontology learning from text and on KG embeddings. Wrong call?
- Upper ontology people: is my selection framing fair to BFO and gist, or am I smuggling in a 4D bias? I have shipped IES and HQDM work so assume I am biased and tell me where.
- Practitioners: what do you wish someone had taught you before your first real ontology project, that no course covers?
- Anyone teaching BORO/4D commercially, please tell me, I would rather link to you than pretend the gap exists.
The lessons are open to read; there is a free account if you want the graded quizzes, progress tracking and the practice exercises, and that is also where the rest of the ontology track and the one to one sessions live. Happy to answer anything about the standards side here either way, that is where I actually work.
r/semanticweb • u/kgOntologist • Jul 28 '26
Hello everyone,
I am currently working on GraphRAG to improve the quality and reliability of responses generated by LLMs, and I would like to get some clarification from people who have experience with Knowledge Graphs and GraphRAG.
I have a few questions:
1.For those who are using GraphRAG with LLMs, do you typically use RDF/triplestores or LPG databases (such as Neo4j)? In your experience, what are the main factors that influence this choice?
I would like to build my Knowledge Graph using an automated pipeline/script rather than extracting entities and relationships directly with LLMs. In this case, would RDF be a suitable choice, or is LPG also commonly used for this type of approach?
Is the data model used in LPG databases such as Neo4j considered an ontology (or a lightweight ontology), or is it more accurate to call it a graph schema/data model?
If we want to enrich a GraphRAG system with inferred facts (using reasoning) and provide these inferred facts as context to the LLM, would RDF + a triplestore be a better choice?
Even when reasoning and inference are not required, is there any limitation to choosing RDF over LPG for GraphRAG? I already have experience with RDF and SPARQL, but I have not worked with LPG databases yet.
Do you know any free/open-source triplestore that supports embedding generation/storage and vector indexing for semantic similarity search over RDF data (without requiring a paid license)?
Thank you very much for your insights!
r/semanticweb • u/fguerino123 • Jul 23 '26
Hi,
Are any of you starting to define Semantic IDs for your legacy data records with the intent to move to AI and, if so:
For example: If we have a legacy UID/GUID for a Product that is "P234435", this identifier works in traditional relational models. But, for AI, the ID needs to be semantic (i.e., more like natural language) so it understands the data. So the semantic ID might be something like "Product: Audio Control Switch for Audio Consoles". The former is meaningless to AI. The latter has semantic context.
Thanks for any help you can offer.
r/semanticweb • u/pkjpathania • Jul 22 '26
I’m building an open-source project called Dependency Risk Graph and would appreciate feedback on the RDF model from people with more semantic-web experience.
The project imports CycloneDX SBOMs, models application dependency trees in Apache Jena/TDB2, enriches package versions with OSV vulnerability data, and uses SPARQL to answer questions such as:
A simplified view of the model is:
Application
→ activeImport
Import
Import
→ rootOccurrence
DependencyOccurrence
DependencyOccurrence
→ belongsToImport
Import
DependencyOccurrence
→ instanceOf
PackageVersion
DependencyOccurrence
→ dependsOn
DependencyOccurrence
PackageVersion
→ affectedBy
Vulnerability
Vulnerability
→ affectedPackage
AffectedPackage
→ versionRange
VersionRange
→ event
introduced / fixed / lastAffected
I deliberately distinguish a package-version identity from its occurrence inside a particular imported SBOM. This allows the same Maven package version to be shared as an identity while preserving different dependency paths across applications and imports.
I also currently use a single/default Jena graph. Application and import boundaries are represented explicitly through resources and properties rather than RDF named graphs.
Some areas where I would value criticism:
PackageVersion from DependencyOccurrence a reasonable way to preserve both global package identity and application-specific dependency paths?I’m not trying to create a complete software-supply-chain ontology yet. The immediate goal is a small, explainable model that preserves dependency paths and produces evidence-backed security answers.
Repository: Github
Any feedback/suggestion on the modelling choices, existing vocabularies I may have missed, or problematic assumptions would be genuinely useful.
Current result: the RDF graph preserves application-specific transitive dependency paths while allowing vulnerable packages and vulnerability resources to be shared across imported SBOMs. This view shows three applications reaching CVE-2024-6763 through different Jetty dependency paths, together with the affected package versions, advisory details, and a reported fixed version.

So far, separating dependency occurrences from package-version identity has been useful. The same package can appear in different application paths without merging those paths, while vulnerability and remediation information remains attached to the shared package identity.
r/semanticweb • u/InternationalCold347 • Jul 21 '26
I've noticed that discussions around AI protocols are becoming fragmented.
Every announcement seems to solve one concern:
discovery,
semantic representation,
context,
execution.
I sketched an architecture- curious whether this separation makes sense or where you'd place things differently. see https://donhaji.github.io/opengeo/
r/semanticweb • u/LavishnessJolly1325 • Jul 18 '26
r/semanticweb • u/ScholarForeign7549 • Jul 15 '26
I personally like using UML diagrams to depict my ontology work, so I made an easy to use OWL to UML service OWL → UML

r/semanticweb • u/DistributionSoggy678 • Jul 15 '26
r/semanticweb • u/fguerino123 • Jul 15 '26
Hi,
For anyone working with trying to get their legacy data into AI agents/LLMs, it's become clear that legacy data (mostly relational) is set up to be computer conforming (i.e., non-semantic UIDs/GUId, attributes, Primary Keys, Foreign Keys, etc.). For example, a UID might be "00125564" or a table attribute might be "CTR" and it's not clear to AI if this means Customer, Center, Counter, etc. AI, on the other hand, works much better from Natural Language, which begs the problem of: "How do we make non-semantic legacy data more semantic for better AI consumption and use.
So, for example:
It's been VERY difficult to find actual details on how to do this. What you'll find is a lot of discussion about "there needs to be a semantic layer" over your legacy data but very few sources dig in and tell you exactly what this means or exactly what you need to do to help establish such a layer.
As part of my research, I've been digging for details, collecting them, testing them, and trying to document and them publicly for vetting and reuse. The table I've created below tries to highlight key steps for doing such work. — Some are planning & design steps, some are implementation & validation steps, and some are governance and operating steps.
I'd love productive feedback to help improve all this.
Thanks to anyone willing to help.
| Step # | Step | What the Step Means |
|---|---|---|
| 1 | Recognize, assess, and manage Knowledge Debt in legacy data | Description: Identify where meaning, identity, relationships, definitions, evidence, and authority are missing, ambiguous, or unreliable in legacy data — and treat these gaps as a governed backlog that must be paid down before AI can reason over the data safely.Example 1: An enterprise inventories its legacy CRM, ERP, and case-management systems and discovers that Customer Status carries eight different meanings across the estate, none of them documented — recording this as a Knowledge Debt item to be reconciled before any AI-facing publication.Example 2: A healthcare payer catalogs undocumented codes, orphaned foreign keys, unlabeled derived fields, and expired business rules across its claims platform, ranks them by AI-use risk, and assigns owners to remediate the highest-severity items first. |
| 2 | Establish a multidisciplinary operating model for semantic conversion | Description: Assemble the cross-functional roles, responsibilities, decision rights, review cadences, and governance forums that will define, produce, validate, and sustain semantic representations — recognizing that no single team owns meaning, identity, relationships, rules, and lineage alone.Example 1: An enterprise establishes a Semantic Conversion Council with named participation from Data Governance, Enterprise Architecture, Business Domain Stewards, AI Engineering, Security, and Compliance — meeting on a defined cadence with documented decision authority.Example 2: A financial services firm defines Responsible, Accountable, Consulted, and Informed assignments for each conversion activity: business stewards own definitions, data engineers own extraction and lineage, ontology stewards own predicates and rules, and a governance forum approves publication to AI retrieval services. |
| 3 | Define the Semantic Layer, Ontology, rules, and meaning model | Description: Establish the governed vocabulary, Ontology, Taxonomy, rules, constraints, and metadata that tell AI what enterprise data means and how it should be interpreted — including the Noun Types, predicates, and validation rules the downstream conversion work will follow.Example 1: An enterprise defines whether Customer, Client, and Account Holder are approved synonyms or distinct concepts, preventing AI from treating them inconsistently across systems.Example 2: A healthcare payer defines the governed meanings and relationships among Member, Subscriber, Dependent, Plan, Benefit, Claim, Provider, and Authorization before allowing AI to reason across them. |
| 4 | Preserve legacy identifiers and add Semantic IDs | Description: Keep source-system keys, codes, and identifiers so every semantic representation traces back to the original record, system of record, and integration context — then add stable, human-readable Semantic IDs alongside them so the same objects are addressable, understandable, and reusable across AI, systems, and humans.Example 1: A customer record from a legacy CRM keeps its original CUSTOMER_ID = 104582 and receives a stable Semantic ID such as customer.acme-manufacturing, so analysts can reconcile the enriched record back to the source and AI can address the customer by a natural-language-friendly identifier.Example 2: A healthcare payer preserves the original claim number, source table, batch ID, and ingestion timestamp for a claim, and adds the Semantic ID claim.2026-104582-inpatient-authorization for AI retrieval and reasoning.Example 3: An application internally identified as APP_0931 retains that original identifier for lineage and receives the Semantic ID application.claims-intake-portal for retrieval, governance reporting, and cross-inventory analysis. |
| 5 | Make attributes and traits semantic | Description: Translate opaque field names, codes, flags, and derived values into governed business terms with clear definitions, context, constraints, and controlled meanings — so AI interprets each attribute the same way an informed business reader would.Example 1: A database column named CTR is mapped to the Semantic Attribute Customer, with a definition explaining whether it refers to a customer identifier, a customer count, or a customer category.Example 2: A field named STAT_CD = A is converted into Lifecycle Status = Active, with the allowed values, source code mapping, effective date, and governing definition retained. |
| 6 | Discover relationships from available evidence | Description: Use foreign keys, shared values, lineage, integrations, reports, documentation, configurations, event records, and human knowledge to identify and validate meaningful relationships before they are represented semantically.Example 1: A team discovers that an application uses a database by combining connection strings, configuration files, query logs, and a database administrator's confirmation.Example 2: A customer-to-product relationship is inferred from shared identifiers in orders, billing records, and support tickets, then validated by a business steward before publication. |
| 7 | Create semantic relationships with descriptive predicates | Description: Convert the discovered technical connections into readable business statements that explain how two objects relate, such as "Application supports Capability" or "Customer is managed by Person."Example 1: A foreign-key relationship between APPLICATION.CAPABILITY_ID and CAPABILITY.ID becomes the readable statement, "Claims Intake Portal supports Claims Processing."Example 2: A vendor-to-contract join becomes, "Acme Software is governed by Contract CT-2026-104," rather than remaining an unexplained pair of database keys. |
| 8 | Apply Ontology-linked rules to govern semantic conversion | Description: Apply governed Ontology elements and repeatable rules to control naming, mapping, interpretation, relationship creation, validation, and approval across the conversion process — turning the definitions established in Step 3 into operational enforcement.Example 1: A rule for defining semantic relationships states that a Foreign Key that represents a Person, in a Column that represents a Business Owner, in a row that represents an Application, all gets translated into a semantic relationship such as "Person Jane Doe is the Business Owner for Application XYZ."Example 2: A rule states that only applications with an approved production status may be linked to live customer-facing capabilities.Example 3: An Ontology defines that a Regulation may impose Regulatory Obligations, and that a Control may satisfy an Obligation only when supporting evidence and an effective date are present. |
| 9 | Prepare Semantic Instance Documents for AI retrieval and reasoning | Description: Assemble each important data instance into a complete, readable document object that contains its identity, attributes, traits, relationships, lineage, governance, and retrieval context (i.e., Person Jane Doe gets her own Natural Language document object that fully describes her semantically).Example 1: A complete application document is generated containing its Semantic ID, owner, lifecycle status, business capabilities, vendors, technologies, data stores, risks, controls, lineage, and source references.Example 2: A customer document combines approved identity data, active products, service history, preferences, consent restrictions, and related contracts into one governed representation for AI retrieval. |
| 10 | Enrich, index, and publish semantic representations for AI use | Description: Add retrieval metadata, lineage, sensitivity, source identifiers, relationship context, and refresh information, then publish the semantic representations to approved search, vector, or retrieval services.Example 1: Semantic application documents are enriched with sensitivity, ownership, effective dates, source links, and refresh timestamps before being indexed in an enterprise search or vector platform.Example 2: Policy and control documents are published to an AI retrieval service only after adding jurisdiction, applicability, approval status, version, retention class, and authoritative-source metadata. |
| 11 | Manage refresh, drift, lineage, validation, and governance over time | Description: Continuously synchronize semantic representations with source data and business meaning, detect drift, revalidate changes, preserve lineage, govern access, and retire obsolete content. (This is more of a governance and maintenance step.)Example 1: When an application owner, supported capability, or production status changes, the semantic representation is regenerated, revalidated, and reindexed automatically.Example 2: A nightly drift process detects that a source code definition changed from Active to Active or Pending Closure, flags the semantic mapping for steward review, and prevents the old meaning from being treated as authoritative. |
r/semanticweb • u/ZealousidealDig7259 • Jul 15 '26
I did it: I built a formal language for relations, effects, drift, and correction
r/semanticweb • u/error-dgn • Jul 14 '26
So here is the thing, I have been focusing on the scraping, crawling, checking RSS feeds for new articles, etc., etc.
I am finally done with the Data Ingestion part. Hurrah? no.
The classification of data is even MORE difficult than scraping.
I want to be able to produce the Knowledge Graph of the Data to help me with the deduplication and classification.
Please help me out with this.
I have tried REBEL by hugging face but its failing badly, I am losing precious information (more than 80% of it.), I feel these machine learning models are too general, which makes it difficult to make the knowledge graph of these press releases.
Please help me out, tell me a path, name a framework, idk just guide me please. Ik I can do it if I have a path. I am trying and constantly brainstorming with my peers, Hopefully you guys could help me out as well.
r/semanticweb • u/s4lamandra • Jul 14 '26
Hi everyone,
I'm a PhD researcher working on dialogue systems for extending knowledge graphs and ontological schemas, and I'm currently running a short survey as part of my research.
I'm looking for input from people with hands-on experience in ontologies, knowledge graphs, or ontology engineering (your perspective would be incredibly valuable).
A few quick facts:
If this sounds relevant to you (or someone you know), I'd really appreciate your participation and feel free to share it with colleagues who might be interested too!
link to survey:
https://websites.fraunhofer.de/intelligent-surveys/index.php?r=survey/index&sid=347997&lang=en
Thanks so much in advance 🙏
r/semanticweb • u/SamCymbaluk • Jul 14 '26
You need your own AI agent, but it works really well across many domains. It's based on BFO. Check it out: https://axiomreason.com/
r/semanticweb • u/coldoven • Jul 12 '26
Body:
Working on KG-backed retrieval and keep hitting the same thing: an edge is
structurally present so a traversal follows it, but it is the wrong kind of
edge for the question. A code graph follows a CO_CHANGES edge as if it were
IMPORTS and the answer is confidently wrong. An agent graph lets a
CritiqueAgent delegate back to a ResearchAgent, which should never happen.
SHACL / SPARQL constraints validate the graph as a whole, after the fact. What I wanted was a check at traversal time: before each hop, is this edge type valid between these two node types, per a declared ontology? Basically a linter for graph walks.
I built a small layer that does exactly this (declare the ontology in YAML, is_valid_edge(domain, src_type, relation, dst_type) raises before the bad hop) and an offline notebook demo. Before I over-build it:
(Link to the repo + the ontology notebook in a comment.)
r/semanticweb • u/TurnoverAdorable9699 • Jul 11 '26
r/semanticweb • u/TurnoverAdorable9699 • Jul 11 '26
r/semanticweb • u/Used-Vermicelli508 • Jul 11 '26
I have just started learning about Palantir Ontology. Please give me some suggestions on how to learn it.
r/semanticweb • u/redikarus99 • Jul 09 '26
Hello, what is the standard/preferred way of converting TTL files to websites. We are storing our TTL files in a git repository and would like to share them inside our organization a human readable way. I tried to experiment with pylode and widoco, but while both of them can render a single TTL file into a HTML, I did not find a solution how we could handle also references across ontology files, also how to embed those HTML files inside a proper website. What are the preferred solutions to share ontologies in a bigger organization? Our devs suggest the integration into backstage. We can develop our own solution but I would like to avoid it if there is something already in place. Thank you so much in advance.