Samuel Douglas Caldwell Jr.

Sam Caldwell

512.712.3095

Sonora, Texas 76950

  • Software engineering
  • Security research
  • DevOps/SRE
  • Cloud Infrastructure

AI-Based Threat-Hunting Tool


Abstract


This specification defines an event-driven threat hunting platform implemented on Amazon Web Services (AWS) that separates telemetry storage and transformation (“Security Lake Layer”) from analyst and model-facing knowledge and reasoning (“Hunter Layer”). The platform ingests heterogeneous threat intelligence and security telemetry into Amazon S3, normalizes and enriches it through AWS Glue, and projects search-optimized views into Amazon OpenSearch Service. Operational failures are isolated via an Amazon SQS dead-letter queue (DLQ) and analyzed by a dedicated Lambda function that emits metrics to Amazon CloudWatch and triggers alerts through Amazon Simple Notification Service (SNS). Hunting logic is represented as structured graphs stored in Amazon Neptune alongside a service catalog/dependency graph, threat models, playbooks, and structured documentation. An LLM hosted on Amazon Bedrock consumes these graphs to propose constrained hunt plans, while AWS Step Functions orchestrates deterministic execution of those plans against OpenSearch and Amazon Athena to produce durable, evidence-backed “Observations” stored in S3 and queryable in Athena. This design aligns with contemporary guidance emphasizing repeatability, evidence preservation, and safety controls for AI-assisted decision support, including prompt/output guardrails and strong validation boundaries (National Institute of Standards and Technology [NIST], 2023; OWASP, 2025).


architectural diagram

Scope


This document specifies:

  1. The logical architecture and data flows across the Data Source Layer, Security Lake Layer, and Hunter Layer.
  2. The graph data models stored in Neptune, including service dependency/support graphs, threat models, and hunt graphs (metadata + indicators).
  3. Deterministic hunt execution via Step Functions and query backends (OpenSearch and Athena).
  4. Operational resilience mechanisms (DLQ, failure analysis, CloudWatch metrics, SNS alerts).
  5. Infrastructure-as-code patterns with Terraform examples for core AWS resources.

Non-Goals


This document does not specify:

  1. Case management (e.g., Jira) is explicitly out of scope.
  2. Version control for hunts/playbooks is external (e.g., Git); Neptune stores structured graph representations used at runtime.
  3. Telemetry events are not written to Neptune; Neptune is exclusively a knowledge/metadata graph store.

Design Principle: Separation of Concerns


The system enforces a strict separation between:

  • Telemetry plane: raw and normalized security data in S3, analytic/search projections in OpenSearch, and query-at-rest via Athena.
  • Knowledge plane: ecosystem context, service relationships, threat models, and hunt definitions stored as graphs in Neptune.
  • Reasoning plane: LLM-based planning and summarization in Bedrock.
  • Control plane: deterministic orchestration and evaluation in Step Functions.

This separation reduces coupling, improves scalability, and prevents “graph-as-SIEM” failure modes in which high-cardinality event streams degrade knowledge stores.


Design Principle: Determinism and Evidence


A core invariant is that the LLM is not the authoritative detector. Instead, the LLM proposes bounded plans; a deterministic evaluator executes queries and computes matches. Observations must carry evidence pointers (e.g., Athena query execution IDs, OpenSearch query metadata, S3 object URIs) to enable replay and audit consistent with incident response practices emphasizing traceability and evidence handling (NIST, 2024).


Design Principle: Safety for AI-Augmented Systems


The platform mitigates known LLM application risks (e.g., prompt injection, unsafe tool use) by using explicit validation, allowlisted actions, and model guardrails. OWASP’s guidance on LLM application risks highlights the need to validate outputs and control downstream actions; NIST AI RMF similarly emphasizes governance and risk controls for AI systems (NIST, 2023; OWASP, 2025).


System Overview


Data Source Layer


This layer includes an arbitrary number of threat intel sources (i.e., data collectors), including but not limited to "active collections" (scraping, polling APIs) as well as passive collection (webhooks, honeypots, log forwarders) and hybrids (e.g. honey token distribution and observation channels.

Data collected at this layer should be structured text (json) or binary (protobuf) data with clear (albeit diverse) schemas for each different feed channel.


Security Lake Layer


The security lake layer consumes the data feed from the data source layer into an s3 bucket, where AWS Glue performs any necessary transformer, normalizer and sanitizer operations, feeding the cleaned results to an OpenSearch service where the data can be queried by the hunter layer.

Invalid data or other data errors are counted by a Lambda function to produce metrics in cloudwatch.


Hunter Layer


The Hunter layer uses AWS Step functions to orchestrate a threat hunt workflow. This workflow sends prompts to AWS bedrock (LLM) which queries the Security Lake and Neptune ecosystem graphs to analyze threat intelligence within the context of the target environment, producing Observations in an S3 bucket.

Threat observations in S3 are queryable by Amazon Athena and are also used to provide thin feedback metric attributes to the Neptune environment graphs. These observations in S3 are individual JSON files Athena can quickly index and consume.


Neptune Graphs


Service Catalog and Dependency Graph


This specification distinguishes composition from runtime dependency to avoid semantic ambiguity.

Node Labels

  • :Product
  • :Service
  • :Component

Composition Edges

  • (Product)-[:HAS_SERVICE]->(Service)
  • (Service)-[:HAS_SUBSERVICE]->(Service)
  • (Service)-[:HAS_COMPONENT]->(Component)

Runtime dependency edge

  • (Product|Service|Component)-[:DEPENDS_ON]->(Product|Service|Component)

Required properties (recommended)

Nodes: stable IDs (product_id, service_id, component_id), name, env, owner_team_id.

Dependency edges: type (e.g., sync_call, async_event, data_read), criticality, since, until, source, confidence.


Service Support Graph


Extends the service graph with organizational support structure and alert.

Node Labels

  • :Team
  • :Person

Edges

  • (Product|Service|Component)-[:OWNED_BY {since, until, primary, source, confidence}]->(Team)
  • (Product|Service|Component)-[:SUPPORTED_BY {since, until, tier, primary, escalation_policy_id, source, confidence}]->(Team)
  • (Team)-[:HAS_MEMBER {since, until, role}]->(Person)
  • (Team)-[:LED_BY {since, until}]->(Person)

This graph supports queries such as “Who owns the services impacted by a dependency chain?” and “Which team is tier-1 support for a targeted component?”


Threat Model Graph


A minimal threat model graph aligned to ATT&CK:

Node labels

  • :Technique (ATT&CK technique IDs)
  • :Detection
  • :DataSource
  • :Control(optional)

Edges

  • (Technique)-[:DETECTED_BY]->(Detection)
  • (Detection)-[:REQUIRES_DATASOURCE]->(DataSource)
  • (Technique)-[:MITIGATED_BY]->(Control) (optional)
  • (Detection)-[:APPLIES_TO]->(Product|Service|Component)(optional)

Hunt Graph


A hunt graph stores hunt metadata and indicators (e.g., IPs, URLs, domains, CVEs) as nodes and edges between them.

  • :Hunt node containing metadata.
  • A set of IOC nodes (:Indicator:*) as graph nodes.
  • Optional logical grouping nodes to encode match semantics without free-form code.

Node labels

  • :Hunt {hunt_id, name, description, severity, confidence, owner_team_id, valid_from, valid_to, commit_sha, source_uri}
  • :Indicator:IP, :Indicator:Domain, :Indicator:URL, :Indicator:Hash, etc.
  • Optional: :IndicatorSet {set_id, semantics: ANY|ALL|K_OF_N, k}
  • Optional: :TemporalWindow {lookback, valid_from, valid_to}

Edges

  • (Hunt)-[:HAS_IOC {role, weight, confidence, source, first_seen, last_seen}]->(Indicator)
  • (Hunt)-[:HAS_SET]->(IndicatorSet)
  • (IndicatorSet)-[:INCLUDES]->(Indicator)
  • (Hunt)-[:TARGETS]->(Product|Service|Component)
  • (Hunt)-[:MAPS_TO]->(Technique)
  • (Hunt)-[:REQUIRES_DATASOURCE]->(DataSource)

Key constraint

Hunt evaluation is deterministic: Step Functions (or a designated evaluator Lambda) interprets IndicatorSet semantics and validates that all plan actions are allowlisted.


Workflows


Ingestion and indexing workflow


  1. Raw data is written to S3 Raw with stable partitions.
  2. _SUCCESS marker is written to the partition prefix.
  3. S3 emits an object-created event to EventBridge.
  4. EventBridge triggers the Ingest Trigger Lambda.
  5. Lambda calls Glue StartJobRun with partition coordinates.
  6. Glue reads raw data, normalizes to clean S3, and updates OpenSearch indices.
  7. Failures are reported to DLQ; Failure Analyzer emits CloudWatch metrics; CloudWatch alarms notify SNS.

Hunt Execution Workflow


  1. Step Functions selects candidate hunts based on:
    • targeted services/products/components
    • ATT&CK techniques relevant to current intel
    • available data sources and time windows
  2. Step Functions fetches hunt and context graphs from Neptune (via a Graph Access Lambda).
  3. Step Functions invokes Bedrock to produce a structured, bounded plan (via LLM Gateway Lambda) under guardrails.
  4. Step Functions validates the plan.
  5. Step Functions executes:
    • OpenSearch queries for fast matching/aggregation
    • Athena queries for deep historical scans (where appropriate)
  6. The evaluator computes match results against hunt graph semantics.
  7. An Observation is produced (structured record + evidence pointers) and written to S3.
  8. Alerts may be published to SNS for high-severity observations.

Feedback Loop


  1. Observations are optionally summarized by the LLM (bounded to evidence).
  2. A Feedback Sanitizer Lambda removes sensitive strings and normalizes feedback.
  3. Sanitized feedback updates either:
    • threat model confidence edges
    • hunt tuning metadata
    • documentation links

Operational Considerations


Scalability


  • Partitioned S3 datasets support parallel ETL and query pruning.
  • OpenSearch scaling must be managed via shard strategy and retention (ISM).
  • Step Functions concurrency should be bounded to protect OpenSearch and Athena from burst overload.

Cost Controls

  • Use ISM for OpenSearch retention.
  • Use S3 lifecycle policies for raw/clean/observation retention tiers.
  • Use Athena workgroups and query limits to bound ad hoc cost exposure.

Audit and Replay

Observations, by design, enable replay: the system stores the hunt identifier and version reference plus evidence pointers. This supports post-incident review and continuous improvement consistent with incident response recommendations emphasizing documentation and continuous refinement (NIST, 2024).


Limitations and Future Enhancements


  • OpenSearch data-plane access: deterministic access should be mediated through a signed-request Lambda to avoid embedding credentials in state machine definitions.
  • Knowledge freshness: service graphs and documentation links require explicit “last_verified/source/confidence” fields to prevent stale context from misleading hunts.
  • Advanced retrieval: structured documentation could be paired with a dedicated search or vector index; this does not change the core separation between telemetry and knowledge planes.
  • Automated coverage analysis: mapping hunts and detections to ATT&CK can support measurable coverage reporting and prioritization (MITRE, n.d.).

References


Amazon Web Services. (n.d.). Creating alarms for dead-letter queues using Amazon CloudWatch. Amazon Simple Queue Service Developer Guide. Retrieved December 23, 2025, from https://docs.aws.amazon.com/AWSSimpleQueueService/latest/SQSDeveloperGuide/dead-letter-queues-alarms-cloudwatch.html

Amazon Web Services. (n.d.). Detect and filter harmful content by using Amazon Bedrock Guardrails. Amazon Bedrock User Guide. Retrieved December 23, 2025, from https://docs.aws.amazon.com/bedrock/latest/userguide/guardrails.html

Amazon Web Services. (n.d.). Index State Management in Amazon OpenSearch Service. Amazon OpenSearch Service Developer Guide. Retrieved December 23, 2025, from https://docs.aws.amazon.com/opensearch-service/latest/developerguide/ism.html

Amazon Web Services. (n.d.). Learning to use AWS service SDK integrations in Step Functions. AWS Step Functions Developer Guide. Retrieved December 23, 2025, from https://docs.aws.amazon.com/step-functions/latest/dg/supported-services-awssdk.html

Amazon Web Services. (n.d.). Neptune graph data model. Amazon Neptune User Guide. Retrieved December 23, 2025, from https://docs.aws.amazon.com/neptune/latest/userguide/feature-overview-data-model.html

Amazon Web Services. (n.d.). Publishing an Amazon SNS message. Amazon Simple Notification Service Developer Guide. Retrieved December 23, 2025, from https://docs.aws.amazon.com/sns/latest/dg/sns-publishing.html

Amazon Web Services. (n.d.). StartJobRun. AWS Glue API Reference. Retrieved December 23, 2025, from https://docs.aws.amazon.com/glue/latest/webapi/API_StartJobRun.html

Amazon Web Services. (n.d.). Using dead-letter queues in Amazon SQS. Amazon Simple Queue Service Developer Guide. Retrieved December 23, 2025, from https://docs.aws.amazon.com/AWSSimpleQueueService/latest/SQSDeveloperGuide/sqs-dead-letter-queues.html

Amazon Web Services. (n.d.). Using EventBridge. Amazon Simple Storage Service User Guide. Retrieved December 23, 2025, from https://docs.aws.amazon.com/AmazonS3/latest/userguide/EventBridge.html

Amazon Web Services. (n.d.). Using Amazon CloudWatch alarms. Amazon CloudWatch User Guide. Retrieved December 23, 2025, from https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/AlarmThatSendsEmail.html

Amazon Web Services. (n.d.). Work with query results and recent queries. Amazon Athena User Guide. Retrieved December 23, 2025, from https://docs.aws.amazon.com/athena/latest/ug/querying.html

HashiCorp. (n.d.). Terraform Registry: hashicorp/aws provider documentation (selected resources). Retrieved December 23, 2025, from https://registry.terraform.io/providers/hashicorp/aws/latest/docs

MITRE. (n.d.). MITRE ATT&CK. Retrieved December 23, 2025, from https://attack.mitre.org/

National Institute of Standards and Technology. (2023). Artificial Intelligence Risk Management Framework (AI RMF 1.0) (NIST AI 100-1). https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-1.pdf

National Institute of Standards and Technology. (2024). Incident response recommendations and considerations for cybersecurity risk management: A CSF 2.0 community profile (NIST SP 800-61r3). https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-61r3.pdf

OWASP. (2025). OWASP Top 10 for Large Language Model Applications. Retrieved December 23, 2025, from https://owasp.org/www-project-top-10-for-large-language-model-applications/