Public job listings imported for 3GIMBALS. These roles were added via admin CSV import or external sources.
Role Overview 3GIMBALS is seeking a Knowledge Graph Engineer to design, build, and maintain the knowledge graph that underpins our unclassified PAI/CAI-based analytic platform. This role is responsible for modeling ontologies, engineering entity resolution and relationship-extraction pipelines, and delivering graph-based analytics that connect people, organizations, locations, events, and other entities across large, heterogeneous data sources. The ideal candidate blends data engineering, semantic modeling, and NLP to turn disparate data into a coherent, queryable graph that powers analytic discovery. Key Responsibilities Ontology & Graph Modeling Design and evolve ontologies, schemas, and taxonomies for entities and relationshipsModel complex, multi-source data into a coherent, queryable knowledge graphDefine and maintain graph data standards, naming conventions, and semantics Entity Resolution & Data Integration Build entity resolution, deduplication, and record-linkage pipelines across disparate sourcesDevelop relationship and event extraction from structured and unstructured data using NLP / information-extraction techniquesIntegrate curated data from the data engineering team into graph ingestion workflowsImplement confidence scoring, provenance, and source attribution for graph assertions Graph Analytics & Query Develop graph queries, traversals, and analytics (centrality, community detection, pathfinding, link analysis)Expose graph capabilities via APIs and query interfaces for analysts and applicationsOptimize graph storage, indexing, and query performance at scale Security & Compliance Ensure the graph and its interfaces meet security requirements for sensitive environmentsImplement access control, encryption, and secure handling of graph dataSupport Authority to Operate (ATO) processes and compliance frameworks Required Qualifications Technical Expertise4+ years of software or data engineering experience, including hands-on knowledge graph workExperience with graph databases (Neo4j, Amazon Neptune, TigerGraph, JanusGraph, or similar)Proficiency with graph query languages (Cypher, Gremlin, or SPARQL)Strong programming skills in Python (Java or Scala a plus)Experience with entity resolution / record-linkage techniques and toolingUnderstanding of ontology and semantic modeling (RDF, OWL, property graphs) NLP & Data Integration Experience with NLP / information extraction (spaCy, Hugging Face, or similar) for entity and relationship extractionExperience integrating heterogeneous structured and unstructured dataFamiliarity with vector embeddings and similarity-based linking Domain Knowledge Experience building or maintaining production knowledge graphsUnderstanding of data provenance, confidence, and source attribution Preferred Qualifications Active security clearance or ability to obtain oneExperience in government, defense, or intelligence contracting environmentsFamiliarity with PAI/CAI data sources and entity-centric analysisExperience with link analysis and network/graph analytics for investigative use casesKnowledge of geospatial-temporal data in a graph contextExperience integrating knowledge graphs with LLM / RAG systems (GraphRAG)Familiarity with federal compliance frameworks (FedRAMP, FISMA, NIST 800-53) Technical Environment Graph: Neo4j / Neptune / JanusGraph; Cypher, Gremlin, SPARQLLanguages: Python (Java/Scala a plus)NLP/ML: spaCy, Hugging Face, embeddings, entity-resolution frameworksData: Integration with platform data pipelines; RDF / property-graph modelsInfrastructure: Docker, Kubernetes, cloud platforms (AWS GovCloud, Azure Government)Security: RBAC, encryption, secure APIs This role owns the connective tissue of the platform: the ontology and graph that let analysts move from isolated records to the relationships, networks, and patterns that drive insight. Salary: $130000 - $180000 per year
Role Overview 3GIMBALS is seeking a Data Engineer to design, build, and maintain the data pipelines and infrastructure that power our unclassified PAI/CAI-based analytic platform. This role is responsible for ingesting, transforming, and curating large volumes of structured and unstructured data from diverse open and commercial sources; building resilient, automated ETL/ELT workflows; and ensuring data is high-quality, well-governed, and analysis-ready for the downstream analytics, knowledge graph, and modeling teams. The ideal candidate is comfortable working with messy, multi-source data at scale within secure development environments. Key Responsibilities Data Pipeline Development & Ingestion Design and build scalable batch and streaming pipelines to ingest structured and unstructured data from PAI/CAI sources, APIs, and third-party feedsDevelop ETL/ELT workflows to normalize, enrich, and transform heterogeneous data into standardized schemasBuild and maintain automated ingestion connectors for web, document, geospatial, and tabular data sourcesManage data orchestration and scheduling using tools such as Airflow, Dagster, or Prefect Data Modeling & Storage Design and maintain data models, schemas, and storage layers across relational, NoSQL, and object storesBuild and maintain data lakes/lakehouses and curated, analysis-ready data martsOptimize partitioning, indexing, and query performance for large datasetsSupport entity resolution and data linking in coordination with the knowledge graph and modeling teams Data Quality, Governance & Lineage Implement data validation, quality checks, and monitoring across pipelinesEstablish data lineage, cataloging, and metadata managementEnforce data governance, provenance tracking, and source attribution appropriate for PAI/CAI dataDocument datasets, schemas, and pipeline logic for downstream consumers Security & Compliance Ensure pipelines and data stores meet security requirements for operation in sensitive environmentsImplement encryption, access control, and secure data-handling practicesSupport Authority to Operate (ATO) processes and compliance frameworks Required Qualifications Technical Expertise 4+ years of data engineering experience building and operating production data pipelinesStrong programming skills in Python and SQL (Scala or Java a plus)Experience with distributed data processing frameworks (Spark, Dask, or similar)Hands-on experience with workflow orchestration tools (Airflow, Dagster, Prefect)Proficiency with relational and NoSQL databases (PostgreSQL, MongoDB, Elasticsearch, etc.)Experience with cloud data platforms and services (AWS, Azure, or GCP) Data & Infrastructure Experience designing data models, warehouses, and lakehouse architecturesFamiliarity with data formats and serialization (Parquet, Avro, JSON, GeoJSON)Understanding of data quality, lineage, and governance practicesExperience with containerization (Docker) and CI/CD for data workflows Domain Knowledge Experience working with large-scale, heterogeneous, or open-source datasetsUnderstanding of data provenance and source-attribution requirements Preferred Qualifications Active security clearance or ability to obtain oneExperience in government, defense, or intelligence contracting environmentsFamiliarity with PAI/CAI (publicly and commercially available information) data sourcesExperience with geospatial data processing (PostGIS, GDAL, or similar)Knowledge of graph data structures and preparing data for knowledge graphsExperience with streaming platforms (Kafka, Kinesis)Familiarity with federal compliance frameworks (FedRAMP, FISMA, NIST 800-53) Technical Environment Languages: Python, SQL (Scala/Java a plus)Processing: Spark, Airflow/Dagster/Prefect, streaming frameworksStorage: PostgreSQL, Elasticsearch, object storage / data lake, ParquetInfrastructure: Docker, Kubernetes, cloud platforms (AWS GovCloud, Azure Government)Security: Encryption at rest and in transit, RBAC, secure data handling This rol