מציאת משרת הייטק בחברות הטובות ביותר מעולם לא הייתה קלה יותר

דרושים Senior Engineer ב-United States, Illinois, Springfield

ממשו את הפוטנציאל שלכם בתעשיית ההייטק עם אקספוינט! חפשו הזדמנויות עבודה בתור Senior Engineer בUnited States, Illinois, Springfield והצטרפו לעוד אלפים שכבר מצאו עבודה בחברות המובילות. התחילו את המסע שלכם עוד היום ומצאו את הקריירה האידיאלית עבורכם בתור Senior Engineer עם אקספוינט.

חברה

אופי המשרה

קטגוריות תפקיד

שם תפקיד (1)

United States

Illinois

Springfield

תאריך יצירה

נמצאו 2 משרות

06.09.2025

Red hat Senior Performance Resilience Engineer - LLM Inference United States, Illinois, Springfield

Limitless High-tech career opportunities - Expoint

שיתוף

התחבר/י כדי להגיש מועמדות

Own the resilience testing roadmap for vLLM and llm-d: define resilience indicators, prioritize fault scenarios, and establish go/no-go gates for releases and CI/CD. Design GPU/accelerator-aware fault experiments that target vLLM...

תיאור:

What you will do:

Own the resilience testing roadmap for vLLM and llm-d: define resilience indicators, prioritize fault scenarios, and establish go/no-go gates for releases and CI/CD
Design GPU/accelerator-aware fault experiments that target vLLM and the stack beneath it (drivers, GPU Operator/DevicePlugin, NCCL/collectives, storage/network paths, NUMA/topology)
Build an automated harness (preferably extending krkn-chaos (https://github.com/krkn-chaos/krkn) ) to run controlled experiments with scoped blast radius, and evidence capture (logs, traces, metrics)
Integrate fault signals into pipelines (GitHub Actions or otherwise) as resilience gates alongside performance gates
Develop detection and diagnostics: dashboards and alerts for pre-fault signals (e.g., vLLM queue depth, GPU throttling, P2P downgrades, KV-cache pressure, allocator fragmentation)
Triage and root-cause resilience regressions from field/customer issues; upstream bugs and fixes to vLLM and llm-d
Explore and experiment with emerging AI technologies relevant to software development and testing, proactively identifying opportunities to incorporate new AI capabilities into existing workflows and tooling.
Publish learnings (internal/external): failure patterns, playbooks, SLO templates, experiment libraries, and reference architectures; present at internal/external forums

What you will bring:

3+ years in reliability, and/or performance engineering on large-scale distributed systems
Expertise in systems‑level software design
Expertise with Kubernetes and modern LLM inference server stack (e.g., vLLM, TensorRT-LLM, TGI)
Observability & forensics skills with experience with Prometheus/Grafana, OpenTelemetry tracing, eBPF/BPFTrace/perf, Nsight Systems, PyTorch Profiler; adept at converting raw signals into actionable narratives.
Fluency in Python (data & ML), strong Bash/Linux skills
Exceptional communication skills - able to translate raw data into customer value and executive narratives
Commitment to open‑source values and upstream collaboration

The following is considered a plus:

Master’s or PhD in Computer Science, AI, or a related field
History of upstream contributions and community leadership, public talks or blogs on resilience, or chaos engineering
Competitive benchmarking and failure characterization at scale.

The salary range for this position is $127,890.00 - $211,180.00. Actual offer will be based on your qualifications.

Pay Transparency

● Comprehensive medical, dental, and vision coverage

● Flexible Spending Account - healthcare and dependent care

● Health Savings Account - high deductible medical plan

● Retirement 401(k) with employer match

● Paid time off and holidays

● Paid parental leave plans for all new parents

● Leave benefits including disability, paid family medical leave, and paid military leave

פרטי המשרה המלאים

משרות נוספות שיכולות לעניין אותך

Red hat Senior Performance Resilience Engineer - LLM Inference United States, District of Columbia, Washington

Red hat Senior Performance Resilience Engineer - LLM Inference United States, Colorado, Denver

Red hat Senior Performance Resilience Engineer - LLM Inference United States, New York, City of Albany

Red hat Senior Performance Resilience Engineer - LLM Inference United States, California, Sacramento

17.04.2025

Red hat Senior Consultant OpenShift Infrastructure Virtualization United States, Illinois, Springfield

שיתוף

התחבר/י כדי להגיש מועמדות

Implement automated, containerized cloud application platform solutions with a focus on infrastructure concerns including networking, storage, virtualization, security, logging, monitoring, and high availability and system resilience. Learn new technologies quickly,...

תיאור:

About The Job

This position requires regular on-site work with clients across North America, so a willingness to travel to customer locations 30-40 weeks per year is required. Applicants must reside within close proximity to a primary airport.

What You Will Do

• Implement automated, containerized cloud application platform solutions with a focus on infrastructure concerns including networking, storage, virtualization, security, logging, monitoring, and high availability and system resilience

• Learn new technologies quickly, including container orchestration, container registries, container build strategies, cloud storage, and software-defined networks

• Travel frequently to work alongside leading financial services, retail, telecommunication, and institutional customers

After joining Red Hat, you will go through an intensive training program on Kubernetes and OpenShift technologies and related DevOps and GitOps topics. Here's how your skill set will evolve and what you'll learn during your first year in the role:

• Understanding of how to build production-ready container and virtualization platforms, integrated with existing enterprise systems

• Knowledge of how to deploy source code into running, scalable containers and virtual machines in automated fashion at enterprise scale

• Practical experience with our offerings

Within 6 months, be ready to implement a routine container platform project by attaining the following:

• Successful, collaborative delivery of customer requirements using Red Hat OpenShift

• Knowledge of how a customer use case can be developed into a project plan and how those requirements align with Red Hat’s technologies

• Understanding of how Red Hat’s technologies can transform software delivery (DevOps/GitOps) practices at large organizations

Within 12 months, begin to demonstrate technical leadership in container platforms by accomplishing the following:

• Successfully implementing complex, large-scale container platform solutions in challenging customer environments

• Helping other peers learn DevOps/GitOps paired with container technologies

• Contributing lessons learned, best practices, and how-tos to our internal and external communities of practice

• Applying new technologies, frameworks, or methodologies to container platforms

What You Will Bring

• Experience leading successful modern cloud platform consulting engagements

• Broad and deep technical experience with VMware ESXi software, including vCenter, VM lifecycle operations using VMware tools

• Logging and alerting functions for SRE operations available from VMware and how they integrate with non-VMware enterprise systems, such as Splunk

• Expert level with VMware VM foundational technologies for networks ([S/DV]Switches/NSX)

• Expert level with VMware VM foundational technologies for storage (Datastores, vSAN, vVols)

• Knowledge of common Add-Ons or third party tools, including Aria/vRealize Suite, SRM, Backup software, Performance tools)

• Experience with technologies including OpenStack, Red Hat Virtualization, Microsoft Hyper-V, Amazon Web Services, and Microsoft Azure a plus

• Experience across one or more vertical industry areas

• Demonstrated track record of working in a strategic advisory role to senior IT and business executives

• Applied knowledge and experience working in agile, scrum, and DevOps teams

• Excellent written, verbal communication and presentation skills

• Willingness to travel to customer locations about 30-40 weeks per year on average across North America

• Degree in computer science or a technical discipline

The salary range for this position is $111,260.00 - $183,580.00. Actual offer will be based on your qualifications.

Pay Transparency

● Comprehensive medical, dental, and vision coverage

● Flexible Spending Account - healthcare and dependent care

● Health Savings Account - high deductible medical plan

● Retirement 401(k) with employer match

● Paid time off and holidays

● Paid parental leave plans for all new parents

● Leave benefits including disability, paid family medical leave, and paid military leave

פרטי המשרה המלאים

משרות נוספות שיכולות לעניין אותך

Red hatSenior Performance Resilience Engineer - LLM Inference

United States, Illinois, Springfield

780558481

06.09.2025

שיתוף

התחבר/י כדי להגיש מועמדות

תיאור:

What you will do:

Own the resilience testing roadmap for vLLM and llm-d: define resilience indicators, prioritize fault scenarios, and establish go/no-go gates for releases and CI/CD
Design GPU/accelerator-aware fault experiments that target vLLM and the stack beneath it (drivers, GPU Operator/DevicePlugin, NCCL/collectives, storage/network paths, NUMA/topology)
Build an automated harness (preferably extending krkn-chaos (https://github.com/krkn-chaos/krkn) ) to run controlled experiments with scoped blast radius, and evidence capture (logs, traces, metrics)
Integrate fault signals into pipelines (GitHub Actions or otherwise) as resilience gates alongside performance gates
Develop detection and diagnostics: dashboards and alerts for pre-fault signals (e.g., vLLM queue depth, GPU throttling, P2P downgrades, KV-cache pressure, allocator fragmentation)
Triage and root-cause resilience regressions from field/customer issues; upstream bugs and fixes to vLLM and llm-d
Explore and experiment with emerging AI technologies relevant to software development and testing, proactively identifying opportunities to incorporate new AI capabilities into existing workflows and tooling.
Publish learnings (internal/external): failure patterns, playbooks, SLO templates, experiment libraries, and reference architectures; present at internal/external forums

What you will bring:

3+ years in reliability, and/or performance engineering on large-scale distributed systems
Expertise in systems‑level software design
Expertise with Kubernetes and modern LLM inference server stack (e.g., vLLM, TensorRT-LLM, TGI)
Observability & forensics skills with experience with Prometheus/Grafana, OpenTelemetry tracing, eBPF/BPFTrace/perf, Nsight Systems, PyTorch Profiler; adept at converting raw signals into actionable narratives.
Fluency in Python (data & ML), strong Bash/Linux skills
Exceptional communication skills - able to translate raw data into customer value and executive narratives
Commitment to open‑source values and upstream collaboration

The following is considered a plus:

Master’s or PhD in Computer Science, AI, or a related field
History of upstream contributions and community leadership, public talks or blogs on resilience, or chaos engineering
Competitive benchmarking and failure characterization at scale.

The salary range for this position is $127,890.00 - $211,180.00. Actual offer will be based on your qualifications.

Pay Transparency

● Comprehensive medical, dental, and vision coverage

● Flexible Spending Account - healthcare and dependent care

● Health Savings Account - high deductible medical plan

● Retirement 401(k) with employer match

● Paid time off and holidays

● Paid parental leave plans for all new parents

● Leave benefits including disability, paid family medical leave, and paid military leave

פרטי המשרה המלאים

משרות נוספות שיכולות לעניין אותך

Red hat Senior Performance Resilience Engineer - LLM Inference United States, District of Columbia, Washington

Red hat Senior Performance Resilience Engineer - LLM Inference United States, Colorado, Denver

Red hat Senior Performance Resilience Engineer - LLM Inference United States, New York, City of Albany

Red hat Senior Performance Resilience Engineer - LLM Inference United States, California, Sacramento

כלי לבניית קורות חיים מקצועיים מבית אקספוינט

הצטרפו למאות שיצרו קורות חיים ושדרגו את הקריירה שלהם

צרו קו"ח

תכננו את מהלך הקריירה הבא שלכם בתעשיית ההייטק עם אקספוינט! הפלטפורמה שלנו מציעה מגוון רחב של משרות Senior Engineer באזור United States, Illinois, Springfield, ומעניקה לכם גישה לחברות הטובות ביותר בתחום. בין אם אתם מחפשים אתגר חדש או שינוי נוף, אקספוינט תקל על מציאת התאמת העבודה המושלמת עבורכם. עם מנוע החיפוש הקל לשימוש שלנו, תוכלו למצוא במהירות הזדמנויות עבודה ולחבור לחברות מובילות. הירשמו היום ועשו את הצעד הבא בקריירת ההיי-טק שלכם עם Expoint.