Basic Concepts

The most important ingestion concepts.

Events and Log Data

The fundamental unit of data in LogScale is an event.

An event is a single, timestamped record that represents something that happened at a specific point in time. For example, a user logging in, a firewall blocking a connection, or an application producing an error.

Events arrive in LogScale as raw log data. This raw data can take many forms depending on the source:

  • Unstructured text - Free-form log lines produced by many applications and operating systems, where the meaning of each part of the line is not explicitly labeled.

  • Semi-structured text - Formats such as Syslog or Common Event Format (CEF), which follow a defined structure but may contain free-form content within that structure.

  • Structured data - Machine-readable formats such as JSON or XML, where each piece of information is explicitly labeled with a key and a value.

Regardless of the format in which data arrives, LogScale uses parsers to interpret and structure it into a consistent set of fields that can be searched and analyzed. See Parsers for more information.

Repositories and Views

In LogScale, data is stored in a repository.

A repository is a dedicated storage container for a collection of events. When you ingest data, you direct it to a specific repository, where it is indexed and made available for search and analysis.

You can have multiple repositories in LogScale, which allows you to organize your data logically. For example, by data source type, business unit, etc.

A view is a virtual 'layer' on top of one or more repositories. Views allow you to query data across multiple repositories as if they were a single data set, but without physically merging or changing the underlying data. Views are useful when you want to:

  • Search across data from multiple sources in a single query.

  • Apply role-based access controls to restrict what data different users or teams can see.

  • Present a filtered or transformed subset of data to a specific audience.

When you are setting up data ingestion, you will need to decide which repository your data should be directed to. This is configured as part of the ingest token setup. See Ingest Tokens for more information.

Ingest Tokens

An ingest token is a credential that authenticates data being sent to LogScale.

Every data source that sends data to LogScale must present a valid ingest token. Without a token, LogScale will not accept the incoming data.

Ingest tokens serve two important purposes:

  • Authentication - The token proves that the sender is authorized to write data to LogScale.

  • Routing - Each token is associated with a specific repository, so LogScale knows where to store the incoming data.

It is recommended practice to create a separate ingest token for each data source. This approach provides several benefits:

  • You can identify which source sent which data, making auditing and troubleshooting easier.

  • You can revoke or rotate the token for one source without affecting any other sources.

  • You can apply different parser assignments to different tokens, so that data from each source is processed appropriately.

Ingest tokens are managed at the repository level within LogScale.

Parsers

A parser is a component in LogScale that interprets raw incoming log data and transforms it into a structured set of fields and values.

Parsers are a critical part of the ingestion pipeline because they determine how your data is organized and what you can do with it once it is stored.

When raw log data arrives at LogScale, it is initially just a string of text. The parser reads that text, identifies meaningful pieces of information within it, and extracts them as named fields. For example, a parser processing a web server access log might extract fields such as:

  • timestamp - When the request was made

  • client_ip - The IP address of the client

  • http_method - The HTTP method used (GET, POST, and so on)

  • status_code - The HTTP response code returned by the server

Once these fields are extracted, they can be searched, filtered, and aggregated in LogScale queries.

LogScale supports several types of parser:

  • Built-in parsers - LogScale includes a set of built-in parsers for common log formats and protocols, such as JSON and Syslog.

  • Package parsers - Many packages available in the LogScale Package Marketplace include parsers specifically designed for a particular data source, such as a firewall vendor or a cloud platform.

  • Custom parsers - If no suitable parser exists for your data source, you can write your own using the LogScale query language and regular expressions.

A parser is assigned to an ingest token, so that all data arriving on that token is automatically processed by the correct parser.

Fields and Tags

Once a parser has processed an incoming event, the data is represented as a collection of fields.

A field is a key-value pair that represents a single piece of information extracted from the event. For example, status_code=404 or username=jsmith.

Fields are the primary way you interact with your data in LogScale. When you write a query, you filter, search, and aggregate based on field names and values.

There are two categories of field in LogScale:

  • Parsed fields - Fields that are extracted from the raw event content by the parser at ingest time.

  • Metadata fields - Fields that LogScale adds automatically, such as @timestamp (the time the event was received) and @rawstring (the original, unmodified log line).

Tags

Tags are a special category of field in LogScale that are used to partition and organize data within a repository.

Unlike regular parsed fields, tags are indexed at storage time and are used by LogScale to determine how data is physically organized on disk.

Because tags influence how data is stored, they have a significant impact on query performance. Filtering on a tag in a query allows LogScale to quickly locate only the relevant segments of data, rather than scanning everything in the repository.

Common examples of fields that are well suited to being used as tags include:

  • The type or source of the log data (for example, #type=firewall).

  • The host or system that generated the event (for example, #host=webserver01).

Tags should be chosen carefully. Because they affect data storage, using too many tags or tags with a very high number of distinct values can have a negative impact on storage efficiency and performance.

Data Collection and Transport

Before data can be ingested by LogScale, it must be collected from its source and transported to the LogScale ingest endpoint.

The method you use to do this depends on the nature of your data source and your environment. LogScale supports several collection and transport mechanisms:

  • Falcon LogScale Collector - A lightweight agent that you install on a host system. The Collector monitors log files and other data sources on that host and forwards events to LogScale in real time. It is well suited to collecting data from servers, workstations, and other host-based sources.

  • Syslog - A widely supported protocol for transmitting log messages over a network. Many network devices, firewalls, routers, and infrastructure components can send log data via Syslog without requiring any additional software. LogScale can receive Syslog data directly or via a Syslog forwarder.

  • HTTP API - LogScale provides an HTTP-based ingest API that allows applications and services to send data programmatically. This is a flexible option for custom integrations and for sources that can make outbound HTTP requests.

  • Integration Packages - Some packages in the LogScale Package Marketplace include pre-configured collection and transport components that handle the ingestion setup for a specific data source automatically.

The choice of collection and transport method does not affect how data is stored or queried in LogScale — it only affects how data gets there. You can use different methods for different data sources within the same LogScale environment.

Packages and the Marketplace

The LogScale Package Marketplace is a library of pre-built integration packages that can simplify and accelerate the process of getting data in from common sources. A package is a self-contained bundle of content that is designed to work with a specific data source or use case.

Packages can include any combination of the following components:

  • Parsers - Pre-built parsers that understand the log format produced by the data source.

  • Dashboards - Pre-built visualizations that present key metrics and activity from the data source in a structured, at-a-glance format. Dashboards can include charts, graphs, tables, and counters that update in real time as new data arrives.

  • Alerts - Pre-configured alert rules for common threat scenarios and operational conditions relevant to that source.

  • Saved Queries - Useful queries that you can run immediately against your data without having to write them from scratch.

Using a package where one is available is generally the fastest way to get value from a new data source, because the parser and supporting content are already configured for you. However, packages are optional. You can always configure ingestion manually if you prefer, or if no package exists for your data source.

Summary

With these concepts in mind, you are ready to move on to the practical steps of getting your data into LogScale. See the Getting Data Into LogScale - Overview for a step-by-step guide to the ingestion process.

Key Ingestion Terms

Understanding the following terms is essential for successfully getting data into LogScale.

Table:

TermDescription
Aggregation Query

A query within the correlate() function that enriches constellations with additional aggregated data—such as counts, sums, or other aggregate values—without affecting whether constellations are found. Aggregation queries run after LogScale identifies a valid constellation. They compute additional values based on events that match their filters and constraints. If an aggregation query matches zero events, the constellation is still emitted. Aggregation queries are distinct from mandatory queries, which define the core pattern required for a constellation to exist.

Backfilling

Process of ingesting historical data into LogScale, with specific guidelines for managing large volumes of historical data efficiently.

Beats

Lightweight data shippers from Elastic (Filebeat, Metricbeat, Winlogbeat). LogScale supports ingesting data from Beats agents through the Elastic Bulk API compatibility layer.

Built-in Parsers

Pre-configured parsers (accesslog, json, syslog, csv, etc.) provided by LogScale for common log formats.

Datasource

An index or segment within a LogScale repository that contains related data. Data sources are created automatically as data is ingested and can be tagged for organization and filtering. They represent logical groupings of data that can be managed independently for retention, access control, and performance optimization.

Elastic Bulk API

An HTTP API protocol compatible with Elasticsearch's bulk ingest format. LogScale provides Elastic Bulk API endpoints to simplify migration from Elasticsearch or integrate with tools that support this protocol.

Event

Single set of fields, originally extracted from a single line of a log file (raw data). An individual event forms the part of any event stream within the LogScale querying system.

Field

A key-value pair extracted from log data that represents a specific attribute or piece of information within a log event.

HEC (HTTP Event Collector)

A widely-supported HTTP API protocol for sending events to LogScale. Originally designed for data migration, HEC is supported by many log shippers and applications. LogScale provides HEC-compatible endpoints for easy integration.

Ingest Rate

The volume of data being ingested per unit time, typically measured in GB/day or events/second. Understanding your ingest rate is critical for capacity planning and licensing.

Ingest Token

Authentication tokens used for secure data ingestion into LogScale repositories.

Integration

A pre-built connection or configuration that enables LogScale to work with external systems, applications, or data sources. Integrations can include log shippers, security tools, cloud services, monitoring systems, and notification platforms. LogScale provides numerous out-of-the-box integrations to simplify data collection and workflow automation.

Log Shipper

A software agent or tool that collects log data from systems and forwards it to LogScale. Log shippers can be lightweight agents installed on servers, containers, or network devices. Popular log shippers include Filebeat, Fluentd, the LogScale Collector, and custom applications using LogScale APIs.

Falcon LogScale Collector

CrowdStrike's lightweight agent for collecting and forwarding logs to LogScale. Falcon LogScale Collector can tail files, collect Windows events, run commands, and enrich data before forwarding. It is the recommended collection agent for most use cases.

Lookup Files

Files stored in a LogScale repository that contain reference data, such as CSV or JSON files mapping field values to additional information. You can use lookup files with the match() function to enrich events during queries, or reference them in parsers to enrich events at ingest time. LogScale supports uploading and managing lookup files through the repository Files page.

Lookup Table

A reference dataset used to enrich events with additional information during querying. Lookup tables can contain mappings like IP addresses to geographic locations, user IDs to names, or error codes to descriptions. They enable data enrichment and correlation without modifying the original log data.

Mandatory Query

A query within the correlate() function that defines the pattern required for a constellation to exist. All mandatory queries must match events for LogScale to identify a constellation. Mandatory queries establish the core pattern that LogScale searches for—for example, a login event followed by a purchase event from the same user. If any mandatory query in the pattern does not find a match, no constellation is created. Mandatory queries are distinct from aggregation queries, which enrich existing constellations but do not affect whether constellations are found.

Metadata

Additional information about events or data that provides context but isn't part of the original log message. Metadata can include source information, parsing details, enrichment data, or system-generated fields. LogScale automatically adds metadata during ingestion and allows custom metadata through parsers and enrichment processes.

Metadata Fields

System-generated fields prefixed with @ (e.g., @timestamp, @rawstring, @id) in LogScale events.

Package

A pre-built bundle of parsers, dashboards, alerts, and saved queries for a specific data source or use case.

Parser

A configuration that defines how to extract structured data from raw log entries. Parsers use regular expressions, built-in functions, or structured formats (JSON, XML, CSV) to identify and extract field values from incoming data. Good parsers are essential for making data searchable and meaningful in LogScale.

Parser Assignment

The mechanism that determines which parser processes incoming events. Parsers can be assigned via ingest token defaults, or LogScale can automatically select parsers based on pattern matching against the raw log text.

Parser Editor

An integrated editing environment in the LogScale web interface for writing and testing parser scripts. The Parser Editor displays the parser script on the left panel and test cases on the right panel, allowing you to run tests against sample log events and inspect the resulting parsed fields.

Real-time

The capability to process and analyze data as it arrives, with minimal delay between data ingestion and availability for searching. LogScale's real-time processing enables live monitoring, immediate alerting, and up-to-the-second dashboards. This is crucial for operational monitoring and incident response.

Repository

Organized collection of data where LogScale stores different sources of data such as events, system logs and metrics.

Syslog

A standard protocol for sending log messages over UDP or TCP. LogScale can receive syslog data directly and parse it using built-in or custom parsers. Syslog is commonly used for network devices, firewalls, and Unix systems.

Tag

A label or marker applied to events, data sources, or repositories for organization and filtering purposes. Tags help categorize data and can be used in queries to focus on specific subsets. Tags can be applied during parsing, through enrichment processes, or manually for organizational purposes.

@timestamp

The @timestamp field is created by LogScale when ingesting data, and represents the timestamp of the original event.

View

A virtual representation that combines data from one or more repositories with optional filtering and access controls. Views enable searching across multiple repositories, restricting access to specific data subsets, or providing simplified interfaces for specific user groups. Views don't store data themselves but provide alternative ways to access repository data.