Engineering integration guide

Datadog + UbiGrowth workflows

Datadog is an observability platform that owns metrics, traces, logs, and alerting for production systems. This guide covers the records that matter, how the connection should be scoped, and what the first bounded workflow should be.

Introduction

Make Datadog part of the workflow, not another silo.

Validate connector availability for your workspace

This guide covers how a team designs a engineering workflow around Datadog with UbiGrowth: which records stay authoritative, how the connection should be scoped, what the first bounded workflow should be, and how to tell whether it worked.

The records that matter are Metric, Monitor, Event, Log, Trace, Dashboard, and Service. Datadog identity is the tag set, not a resource id. Two hosts reporting the same service tag are one service; a tag typo creates a second service that looks real and silently splits every metric.

Datadog is not currently on UbiVibe's verified connector list. This page is an implementation design reference: use it to specify the workflow, then validate whether the connection is available and correctly scoped for your workspace before you make it a dependency. The verified UbiVibe connections today are Salesforce, HubSpot, Gmail, Google Drive, Slack, and GitHub.

The platform layer is the usual destination for this connection, because the value shows up as governed context and execution shared across more than one team.

Why teams evaluate this connection

Integrations create value when they remove operating friction.

The first design decision is not which API endpoint to call; it is which system owns the record, what event should trigger work, who owns the exception path, and what successful completion means.

Engineering systems produce more signal than any other part of the business and the least usable summary. Datadog knows exactly what happened; turning that into something the rest of the company can act on is manual work that nobody owns.

The second problem is direction of risk. An integration that reads engineering activity is low-risk and useful. An integration that can act on infrastructure or production systems is a different category entirely, and the two are often discussed as if they were the same project.

You're likely here because

  • Engineering activity is invisible outside the engineering team
  • Incident context has to be reassembled manually every time
  • Internal tool requests sit behind product work indefinitely

Record model

What a Datadog integration actually reads and writes.

Integration design starts from the objects the system really exposes, not from a generic connector diagram. These are Datadog's.

MetricMonitorEventLogTraceDashboardService

Identity and matching

Datadog identity is the tag set, not a resource id. Two hosts reporting the same service tag are one service; a tag typo creates a second service that looks real and silently splits every metric.

Start here

Read Monitor state transitions for one service and correlate alerts to deploys, so recurring noise becomes visible as a pattern rather than a page.

What this will not do

It will not tell you what broke. Datadog reports symptoms; the causal step needs the deploy and change history alongside it.

The constraint to plan around

Metric queries are billed and rate-limited, and high-cardinality tags multiply custom metric cost. A dashboard query run on a schedule can become a material line item.

Build notes

What you actually have to reason about in Datadog.

The fields that carry meaning, how the connection authenticates, and whether the event surface can be trusted. This is the part that decides whether the integration works in month three.

FieldWhy it matters
tagsidentity is the tag set — service, env, version — and a typo silently forks a service in two
monitor state / transitionthe alert-worthy event is the transition, not the current state
custom metric cardinalitythe cost driver, since each unique tag combination is a billable custom metric
span / trace idthe correlation keys between traces, logs, and metrics
aggregation and rollupa query over a long window silently rolls up, changing what the number means

Authentication

API key for submission and an application key for queries, and the two are not interchangeable. Application keys inherit the creating user's scopes, so a key made by an admin grants admin-level reads to whatever holds it.

Events and delivery

Webhooks fire from monitor notifications rather than as a general change feed. The Events API accepts inbound deploy and change markers, which is what makes alert-to-deploy correlation possible.

Workflow

How the Datadog workflow runs.

The operating sequence, from reading the source system through to the result landing back where it belongs.

01Connect read-only first02Assemble the engineering picture03Build the internal surface04Bound any action path

Step 01

Connect read-only first

Datadog is connected with scoped, read-only credentials so context and reporting value can be proven without any action risk.

Step 02

Assemble the engineering picture

Delivery, incident, or operational activity is summarized in a form the rest of the business can act on rather than a raw feed.

Step 03

Build the internal surface

Launch produces the dashboard or internal tool that was never going to clear the product backlog, reviewed like any other internal service.

Step 04

Bound any action path

If the workflow needs to act, that path is specified separately with explicit scope, approval, and a record of what ran.

Design decisions

The engineering decisions this connection forces.

Each of these has to be settled before the Datadog workflow is allowed to write anything.

01Separate read from act02Make execution paths explicit

Step 01

Separate read from act

Reading Datadog for context and reporting is a different risk decision from letting a workflow act on it. Do not bundle them into one project.

Step 02

Make execution paths explicit

Any action that reaches a real environment should run through a reviewable execution path with a record of what ran, not an implicit side effect.

Implementation path

How to implement the Datadog workflow.

  1. 01

    Start read-only against Datadog and produce something the team already wants: delivery visibility, incident context, or an operational summary.

  2. 02

    Use scoped credentials rather than a shared token, and confirm what the scope can actually reach.

  3. 03

    Build the internal surface in Launch, and review the result as you would any other contribution.

  4. 04

    After alert-to-deploy correlation works, retire the alerts that never precede a real incident, which is the change that makes on-call sustainable.

Governance

Controls that matter.

01

Control 01

Credentials are scoped and workspace-approved; no shared secret belongs in a prompt or in generated code.

02

Control 02

Actions that reach production systems run through explicit, reviewable execution paths.

03

Control 03

Generated code and configuration are reviewed on the same terms as any other change.

Failure modes

How a Datadog integration breaks in production.

Not generic integration advice. These follow from how this system actually behaves, which is why they look nothing like the list on the next guide over.

Symptom 01

A service appears twice with half its traffic each.

Cause

A tag typo forked the identity — Datadog identity is the tag set.

Fix

Validate tags at instrumentation time and alert on new unexpected service tag values.

Symptom 02

The monthly bill rises sharply with no traffic change.

Cause

High-cardinality tags multiplied custom metric count.

Fix

Audit tag cardinality before shipping, and never tag by user id or request id.

Symptom 03

A query over a quarter shows different values than over a week.

Cause

Long windows roll up, changing the aggregation.

Fix

Specify the rollup explicitly so the number means the same thing at every window.

What changes at scale

Query and ingestion are billed separately, and dashboards on auto-refresh are a recurring cost. Cache query results for anything displayed continuously.

Examples

What a working Datadog workflow looks like.

Bounded scenarios rather than a feature list. Each one can be verified against work the team already does.

Observability workflows

With Datadog connected read-only, delivery and operational activity can appear alongside commercial context instead of living in a separate report.

Internal tool that was stuck in the backlog

A small tool reading Datadog gets built in Launch and reviewed like any other internal service, without consuming sprint capacity.

Limitations and considerations

What to validate before you depend on this.

  • Metric queries are billed and rate-limited, and high-cardinality tags multiply custom metric cost. A dashboard query run on a schedule can become a material line item.
  • Submitting high-cardinality custom metrics is a billing event. An integration that tags by user id or request id can add a material monthly cost before anyone notices, and the data is already ingested.
  • When there is no on-call process to receive what you find. Observability integrations create alerts, and alerts with no owner become noise that trains people to ignore the tool.
  • Write or action access to Datadog is a materially different risk decision from read access and should be scoped, reviewed, and approved separately.
  • Generated code and configuration still require review. Speed of production does not change ownership of what ships.

FAQ

Datadog integration questions.

What records does a Datadog integration actually work with?

The primary records are Metric, Monitor, Event, Log, Trace, Dashboard, and Service. Datadog identity is the tag set, not a resource id. Two hosts reporting the same service tag are one service; a tag typo creates a second service that looks real and silently splits every metric.

What should the first Datadog workflow be?

Read Monitor state transitions for one service and correlate alerts to deploys, so recurring noise becomes visible as a pattern rather than a page.

What will a Datadog integration not do?

It will not tell you what broke. Datadog reports symptoms; the causal step needs the deploy and change history alongside it.

What is the main constraint to plan around?

Metric queries are billed and rate-limited, and high-cardinality tags multiply custom metric cost. A dashboard query run on a schedule can become a material line item.

What changes about a Datadog integration at scale?

Query and ingestion are billed separately, and dashboards on auto-refresh are a recurring cost. Cache query results for anything displayed continuously.

How does authentication work for Datadog?

API key for submission and an application key for queries, and the two are not interchangeable. Application keys inherit the creating user's scopes, so a key made by an admin grants admin-level reads to whatever holds it.

Does Datadog support webhooks, and can they be trusted?

Webhooks fire from monitor notifications rather than as a general change feed. The Events API accepts inbound deploy and change markers, which is what makes alert-to-deploy correlation possible.

What is the risk of writing to Datadog?

Submitting high-cardinality custom metrics is a billing event. An integration that tags by user id or request id can add a material monthly cost before anyone notices, and the data is already ingested.

When is connecting Datadog the wrong call?

When there is no on-call process to receive what you find. Observability integrations create alerts, and alerts with no owner become noise that trains people to ignore the tool.

What should a Datadog integration automate first?

Start with one bounded workflow that removes a measurable handoff, duplicate-entry step, reporting delay, or follow-up gap. Expand only after the first workflow is reliable.

Does UbiGrowth require Datadog to be replaced?

No. The operating model is designed around connecting to systems that should remain authoritative and building workflows around them rather than forcing a wholesale replacement.

Is connector availability identical for every workspace?

No. Availability can depend on provider configuration, authentication, scopes, workspace setup, and deployment state. Validate the required connection before treating it as an operational dependency.

Can the workflow act on Datadog, not just read it?

Action paths are possible but should be treated as a separate, bounded project with scoped credentials, explicit approval, and a record of what ran.

How are credentials handled?

Through workspace-approved, scoped grants. A shared secret pasted into a prompt or embedded in generated code is not an acceptable pattern.

What is a safe first integration?

A read-only workflow that produces something the team already wants from Datadog — delivery visibility or incident context — before any action path is considered.

How this access is governed

What ARIA is allowed to do in Datadog, and who decides.

Connecting Datadog is a permission decision, not just a setup step. These are the controls that decide what ARIA can reach, what it can change, what gets recorded, and how you take the access back.

Required permissions

ARIA works through the scopes the connection was granted, and no others. Authorization happens at the provider, so the permissions being requested are shown by the system itself before anything is connected.

What it can reach

Reachable systems are the intersection of what your organization approved in the connector registry and what the requesting identity is permitted to use. Identity resolves before execution, not after.

What it can do

Actions run through explicit execution paths with state, spend, and failure boundaries — a bounded worker path rather than an open-ended agent loop with a credential.

Credential handling

Credentials live in the governed connection layer and are resolved through canonical connection identity. They are not pasted into individual workflows, prompts, or generated artifacts.

Action logging

Execution carries state and traces: what triggered the work, which connection it used, and what came back — including an explicit failure when something did not run.

Approval and revocation

Consequential actions can be made to require a person to approve them. Access can be changed or revoked at the connection, and ARIA loses that reach without unpicking the work already completed.

Start with ARIA

Ask ARIA to run this integration.

Describe the outcome you need across this system. ARIA works out the scopes, data, and actions the job requires, and operates inside the access you grant — which you can change or revoke.

  • ARIA acts only through the systems and permissions you connect.
  • Connections use scoped credentials you can change or revoke.
  • Actions are recorded, and consequential ones can require approval.

Goes to UbiGrowth, with the page you asked from attached. We do not sell or share it. Prefer to talk? Call 972-823-1294.

Start here

Turn the integration into a working business outcome.

Start with ARIA to describe the outcome, then continue into the product path that fits the workflow. Connector availability and required scopes should be validated for the specific workspace before production use.