All positions
Careers

Data Engineer — Analytics & Data Lake

Remote (U.S. clients, LatAm-based)

  • ELT
  • Cloud Data Warehouse
  • BI & Analytics

About the Role

We’re looking for a Data Engineer to design, build, and maintain the data pipelines and analytics layer that power reporting for the customers we serve — whether that’s a single dedicated customer or several at once. You’ll manage the ingestion of data from source systems (e.g., e-commerce, CRM, and marketing platforms) into a central cloud data warehouse using modern data-integration tools, then model clean, well-structured data marts on top to support downstream reporting and analytics.

This is a hands-on, end-to-end role: you own pipelines, data models, BI, and data quality across each customer environment. You’ll build and maintain the transformation logic that turns raw data into curated, business-friendly datasets — consistent, accurate, and performant — and work closely with analysts and business stakeholders to translate reporting requirements into scalable models that power BI dashboards. You’ll also monitor pipeline health, troubleshoot sync and transformation issues, and keep documentation current across every customer instance. Work is triaged and ticketed through a delivery PM, so you stay focused on building — not fielding ad-hoc requests.

What You’ll Own

Data Lake & Pipelines

  • Own the ELT process (e.g., Airbyte) that syncs e-commerce platform data (e.g., Shopify) into a relational analytics warehouse (e.g., Postgres); keep syncs healthy, monitored, and documented.
  • Extend ingestion to new sources as reporting needs grow — email/subscriber marketing platforms, web analytics platforms, and e-commerce admin APIs.
  • Design pipelines for parallel, fault-tolerant collection — sources fetched concurrently, with graceful degradation when one feed is down, so a single failed source never blocks a report.
  • Model and maintain the analytics schema: clean, documented tables that both BI and the AI layer can query reliably.

Analytics & BI

  • Build and maintain BI dashboards and models on the analytics warehouse.
  • Run data validations before every sign-off — reconcile warehouse numbers against the system of record so commission, order-volume, and revenue reporting can be trusted.
  • Build a DTC marketing-analytics dashboard: unify commerce, web analytics, and ad-platform data into campaign performance, conversion funnels, and drop-off tracking — accounting for measurement discrepancies across systems and multi-site attribution.

AI-Adjacent Data Enablement

  • Build and maintain a conversational-AI-to-BI connection (e.g., via an AI/agent integration protocol) so leadership can query analytics conversationally.
  • Provision clean, well-modeled data feeds for the customer’s AI agent platform — reporting and data-monitoring agents that read the warehouse for scheduled reports and pipeline-health alerts.
  • Treat BI as your warehouse-visualization tool; business users consume insights through the AI layer, not the BI tool directly.

Data Quality, Governance & Compliance

  • Hold the order-data-only line: this stack ingests order/commerce data, not customer PII/PHI, without explicit approval — protect that boundary and escalate requests that cross it.
  • Enforce data-separation and privacy rules; restricted contact information is never exposed in reporting.
  • Never store or log sensitive data outside an approved audit trail; escalate any regulatory or clinical data question rather than guessing.
  • Document pipelines, schemas, and runbooks so the work is transferable and auditable.

How You Work

  • Intake through the delivery PM. Data requests route through a dedicated Slack channel; the PM triages and creates tickets (JIRA). You work a prioritized, tracked, visible queue — not one-off DMs.
  • Embedded with each customer team. Your day-to-day partners are the customer’s operations and revenue stakeholders; you coordinate with other engineers on shared source data.
  • Delivery standards. Work is reviewed against a managed-service quality bar, and validations always precede customer sign-off.

Technical Requirements

Required

  • 3–6 years in data engineering / analytics engineering, shipping production ELT and BI independently.
  • Strong SQL and Postgres data modeling (analytics/warehouse schemas, not just application CRUD).
  • Hands-on with a modern ELT tool — Airbyte strongly preferred (Fivetran / dbt / Meltano transferable).
  • BI development experience — Metabase or Tableau strongly preferred (Looker / Power BI transferable).
  • Comfortable integrating e-commerce and marketing APIs — e.g., Shopify, Klaviyo, GA4, ad platforms.
  • Practical data-quality instincts: validation, reconciliation, and graceful handling of messy source data.
  • Clear written communication in English — you document what you build and flag data caveats plainly.

Nice to Have

  • E-commerce / DTC analytics background (Shopify ecosystem a plus).
  • Exposure to LLM/AI data plumbing — MCP connectors, RAG data prep, or feeding structured data to agents.
  • Experience in a regulated / health-adjacent data environment (PHI awareness, audit trails).
  • Operational familiarity with Railway / Fly.io / AWS RDS.

AI-First Data Engineering Culture

At NearShift, AI isn’t a side project — it’s how we work. You’ll wire analytics directly into LLM-powered agents, use AI tools to move faster, and build the data foundations that make conversational analytics reliable. We look for engineers who are curious about where data engineering and AI meet.

Core Competencies

  • Analytical thinking and a results-driven mindset.
  • Ownership — you take a pipeline or dashboard from raw source to trusted output.
  • Practical rigor around data quality, validation, and reconciliation.
  • Clear documentation and plainly-flagged caveats.
  • Collaboration and knowledge sharing across product, operations, and revenue teams.
  • Proactivity, continuous improvement, and AI-first thinking.

What Makes You a Great Fit

You care about data people can trust. You know a dashboard is only as good as the pipeline and the validation behind it, and you’d rather catch a discrepancy before sign-off than explain it afterward. You enjoy:

  • Owning a data stack end-to-end — pipelines, models, BI, and quality.
  • Turning raw, messy data into reporting leadership actually uses.
  • Building the data foundations that let AI agents answer questions reliably.
  • Working as an embedded member of a customer team — one or several — across full U.S. time-zone overlap.
  • Documenting your work so it’s transferable and auditable.

What We Offer

Compensation & Benefits

  • Competitive salary based on your experience and skills.
  • Fully remote work with full time-zone overlap with U.S. clients.
  • Professional development budget and dedicated time for skill development and self-study.
  • Access to cutting-edge AI tools and platforms.

Career Growth

  • Long-term placement with leading North American clients.
  • International exposure and direct partnership with U.S. business and engineering leaders.
  • The opportunity to shape customers’ data platforms and influence technical decisions.
  • A path to grow into senior / lead data engineering and AI-data roles.

About NearShift

NearShift places AI-first engineering talent from Latin America with leading North American companies — vetted, embedded, and ready to ship. We pair strong engineers with U.S. clients across full time-zone overlap, so our people work as true members of the client’s team, not an offshore vendor at arm’s length. We care about clean pipelines, data people can trust, and engineers who keep getting better at their craft.

Ready to own a real data stack?

If you’re excited about clean pipelines, trustworthy analytics, and the place where data engineering meets AI — let’s talk.

Apply for this role
Apply for this role