Business

Databricks launches native IP functions for network data

Databricks has made its native IP functions generally available on Runtime 18.3+, allowing enterprises to run high-performance network analytics directly within their data lakehouse.

Databricks AI1 day agoBusiness
Image: Databricks AI

Databricks has announced the general availability of native IP functions on Databricks Runtime 18.3 and above. These built-in SQL, PySpark, and Scala functions are designed to parse, validate, canonicalize, and join IPv4 and IPv6 addresses alongside CIDR blocks. Optimized within the Photon engine, these tools allow demanding network workloads to run in seconds rather than hours. In head-to-head benchmarks against a leading cloud data warehouse, Databricks completed IP CIDR joins up to 3.1x faster and between 2x and 6.4x cheaper.

The new toolkit includes specialized operations such as ip_cidr_contains to test if an address falls within a block, ip_host to normalize addresses, and ip_version to identify protocol versions. It also features ip_prefix_length, ip_network, ip_network_first, ip_network_last, and ip_cidr. To optimize performance, practitioners can use ip_as_binary and ip_as_string to convert and store compact representations. Additionally, error-tolerant variants like try_ip_host, try_ip_cidr, try_ip_as_binary, and try_ip_as_string return null values instead of failing when encountering malformed rows in massive datasets.

For data engineers and security analysts, this release eliminates the need for brittle regex parsing, slow user-defined functions, and complex bitwise math. Previously, analyzing high-volume network logs required specialized stacks or pre-expanding CIDR blocks, which often excluded modern traffic or created governance silos. Now, practitioners can run threat detection, fraud investigation, and network observability queries directly in the lakehouse using standard SQL.

Integrators are already deploying these capabilities for large-scale enterprise workloads. Rearc, a consultancy helping firms build cloud and AI platforms, implemented these native functions for a major financial client processing more than 30TB of data per day. By replacing ad-hoc implementations with native SQL operations, they achieved simpler, more maintainable pipelines without compromising on processing speed.

This is our own summary of reporting by Databricks AI

More in Business