Job Description
Insight Global is looking for a Data Engineer, GenAI Data Systems for one of our largest customers in the Bay Area who will hit the ground running on Data Engineering work, and someone who can contribute broadly across a variety of operational engineering tasks, not just traditional data pipeline development.
This role can pay between $75-$105 depending on years of experience and skillset.
We are a company committed to creating diverse and inclusive environments where people can bring their full, authentic selves to work every day. We are an equal opportunity/affirmative action employer that believes everyone matters. Qualified candidates will receive consideration for employment regardless of their race, color, ethnicity, religion, sex (including pregnancy), sexual orientation, gender identity and expression, marital status, national origin, ancestry, genetic factors, age, disability, protected veteran status, military or uniformed service member status, or any other status or characteristic protected by applicable laws, regulations, and ordinances. If you need assistance and/or a reasonable accommodation due to a disability during the application or recruiting process, please send a request to HR@insightglobal.com.To learn more about how we collect, keep, and process your private information, please review Insight Global's Workforce Privacy Policy: https://insightglobal.com/workforce-privacy-policy/.
Required Skills & Experience
- Bachelor's Degree in Computer Science or related engineering field, or equivalent experience.
- 4+ years of professional experience in data engineering or a closely related field.
- Experience designing, building, and operating production ETL/ELT pipelines for large, complex datasets.
- Strong SQL and Python proficiency for data pipelines, automation, and production software.
- Familiarity with CI/CD, testing frameworks, and version control (e.g., Git).
- Experience with workflow orchestration (e.g., Airflow, Dagster, or Databricks Jobs/Lakeflow).
- Demonstrated experience designing and maintaining data models, data pipelines, and databases (relational and/or lakehouse/warehouse).
- Experience designing and operating high-volume logging or telemetry systems, with a focus on efficiency and monitoring.
- Experience building and maintaining dashboards, reports, and alerting, and keeping metric definitions accurate over time.
- Hands-on experience with Databricks, Spark, Delta Lake, or similar lakehouse/big-data platforms.
- Experience with a major cloud platform (AWS, Azure, or GCP) and large-scale storage systems.
- Proficiency using AI tools for coding and data tasks, paired with strong first-principles reasoning about design, coding, and datasets
Nice to Have Skills & Experience
- 8+ years of experience in data engineering or analytics engineering.
- Experience building data pipelines specifically for annotation projects running in annotation platforms (human-in-the-loop / labeling workflows).
- Direct experience with SuperAnnotate.
- Experience supporting LLM/VLM or other ML model training data pipelines.
- Experience working with large-scale, high-volume, multi-modal datasets (e.g., text, image, video, audio), with attention to storage and processing efficiency.
- Front-end or full-stack experience (JavaScript/TypeScript/HTML) for internal tooling and lightweight data/annotation UIs.
Benefit packages for this role will start on the 1st day of employment and include medical, dental, and vision insurance, as well as HSA, FSA, and DCFSA account options, and 401k retirement account access with employer matching. Employees in this role are also entitled to paid sick leave and/or other paid time off as provided by applicable law.