Turing

Remote site reliability engineer jobs

We, at Turing, are looking for site reliability engineers who will be responsible for automating solutions including capacity and performance planning, managing risks, disaster response, and on-call monitoring. Here’s your chance to work with elite U.S. companies and collaborate with top professionals across the globe.

Find remote software jobs with hundreds of Turing clients

Job description

Job responsibilities

  • Build software applications to help operations and support teams
  • Gather and analyze metrics to help in performance tuning and troubleshooting errors
  • Contribute to system design consulting, platform management, and capacity planning
  • Develop sustainable systems and services with automation and uplifts
  • Improve feature development speed and system reliability through optimization of on-call processes
  • Prepare documentation of historical knowledge concerning software development, support, IT operations, and on-call duties
  • Monitor application performance and keep the sites up and running

Minimum requirements

  • Bachelor’s/Master’s degree in Engineering, Computer Science, or IT (or equivalent experience)
  • At least 3+ years of experience as a site reliability engineer (rare exceptions for highly skilled engineers)
  • Proficient understanding of operating systems (Linux/Windows)
  • Expert knowledge of DevOps concepts and best practices
  • Expertise in CI/CD implementation
  • Hands-on experience in troubleshooting issues
  • Knowledge of one or more high-level programming languages like Python, Java, JavaScript, C/C++, Ruby, etc.
  • Experience with distributed storage technologies and dynamic resource management frameworks
  • Fluent in English to communicate effectively
  • Ability to work full-time (40 hours/week) with a 4 hour overlap with US time zones

Preferred skills

  • Working knowledge of code versioning tools such as Git
  • Proactivity in finding issues, performance bottlenecks, and areas for improvement
  • Passion for automation, coding skills, and software-centric mindset
  • Understanding of distributed computing, cloud-native applications, application monitoring, and database management
  • Excellent organizational and interpersonal skills

Interested in this job?

Apply to Turing today.

Apply now

Why join Turing?

Elite US Jobs

1Elite US Jobs

Turing’s developers earn better than market pay in most countries, working with top US companies.
Career Growth

2Career Growth

Grow rapidly by working on challenging technical and business problems on the latest technologies.
Developer success support

3Developer success support

While matched, enjoy 24/7 developer success support.

Developers Turing

Read Turing.com reviews from developers across the world and learn what it’s like working with top U.S. companies.
4.65OUT OF 5
based on developer reviews as of June 2024
View all reviews

How to become a Turing developer?

Work with the best software companies in just 4 easy steps
  1. Create your profile

    Fill in your basic details - Name, location, skills, salary, & experience.

  2. Take our tests and interviews

    Solve questions and appear for technical interview.

  3. Receive job offers

    Get matched with the best US and Silicon Valley companies.

  4. Start working on your dream job

    Once you join Turing, you’ll never have to apply for another job.

cover

How to become a Site Reliability engineer ?

As software development became faster and more complex, traditional software teams had trouble keeping up. To help with the transition of workflows from development to production applications, they introduced DevOps.

However, it became increasingly apparent that this system needed greater reliability and performance in order to stay competitive. This is where the field of site reliability engineering comes into play.

Site reliability engineering blends software engineering practices with information technology (IT) engineering practices to create highly reliable systems. Site reliability engineers are responsible for ensuring the reliability of all aspects of the full stack, from the front-end, customer-facing applications all the way through to the database and hardware infrastructure.

What is the scope in Site Reliability engineering?

The role of SRE (Systems and Release Engineer) is ideal for assessing the newest development in the DevOps world, expanding your knowledge and skills in high-demand areas such as infrastructure automation, release engineering, and continuous delivery. As an SRE, you’ll be highly creative, stimulated, and technically challenged every day.

Site reliability engineers are crucial to most organizations. These professionals are in high demand at successful tech companies that have large data centers and complex technical challenges. They can also be inspirational from both a financial and workplace culture perspective. Google considers them scarce resources.

What are the roles and responsibilities of a Site Reliability engineer?

Site reliability engineering (SRE) refers to software engineering approaches used by organizations to manage their IT operations. SRE teams use software tools as a way to automate operations and solve problems in a timely manner.
Software reliability engineers (SREs) are software engineers who have Unix systems administration, networking, and software engineering experience. SREs also have polished programming skills because they regularly use automation to reduce human labor and increase reliability.
Software Release Engineering (SRE) transfers the tedious work traditionally done by DevOps and operations teams to software engineers who can use automation and software to optimize processes.
Site reliability engineers spend half their time doing development work, and the other half doing operations duties, such as responding to outages and incidents and being on call.

The roles and responsibilities of a site reliability engineer include

  • Building software to help Operations and Support Teams
  • Conducting Post-Incident Reviews
  • Documenting the knowledge to ensure a seamless flow of information between teams
  • Implementing strategies to increase system reliability and performance through on-call rotation
  • Fix cases related to support escalation
  • Incorporate various software engineering aspects to develop and implement services that improve IT and support teams
  • Optimize the Software Development Life Cycle (SDLC) to boost service reliability

How to become a Site Reliability engineer?

You can become a site reliability engineer in the following ways:

  1. Bachelor's degree: It is mandatory for the developer to have a Bachelor’s degree or Master’s degree. This helps with growth in the software field and also aids in easy understanding of technical aspects of the job.
  2. 2+ years experience in operations or software engineering role: It helps if you have some previous experience working as a software engineer. This will give you an advantage over other candidates while trying for SRE positions.
  3. Required skills: You must have the following technical skills.
  • Experience with cloud-continuous deployment based software development lifecycles
  • Expertise in infrastructure automation technologies

Along with technical skills you must have a strong foundation of non-technical skills as well. What you need:

  • Excellent verbal and written communication skills
  • Strong problem-solving skills
  • Passion and curiosity for technology
  • Keenness to provide support for teams or customers.

Now let us discuss the skills and methods you will need to learn to become a successful site reliability engineer:

Interested in remote Site Reliability jobs?

Become a Turing developer!

Apply now

Skills required to become a Site Reliability engineer

Fundamental skills are important in helping you land high-paying site reliability engineer jobs. Here is what you need to know!

1. DevOps

DevOps refers to a set of practices that promote better collaboration and widespread automation of the processes happening between operational and development teams. It can be extended to other business units as well.

DevOps is a new cultural movement combining software development, operations, and engineering. It stimulates the adoption of agile practices that are continuous in nature and enable continuous delivery of small batches to customers.

2. Python

Python is easy to learn. It is a high-level, dynamic language with an interpreted structure to make debugging errors relatively painless. Which helps programmers rapidly develop working application prototypes. This feature has earned Python a reputation as a language well-suited for coding. Because Python supports cross-platform operating systems, it is a good choice for programmers. Especially those who do not want to spend time writing separate programs for different operating systems.

3. Go

Go was created for applications relating to network infrastructure and was intended to replace Java and C++. It is used in cloud-based or server-side (web) applications. With DevOps, site reliability automation, micro-controller programming, robotics, and games also common users of Go. Go is also used in the world of artificial intelligence and data science.

4. CI/CD

Continuous integration/continuous delivery (CI/CD) is a software development process in which code is automatically built and tested as new code is added. CI/CD can improve the effectiveness of a software team by reducing the risk of errors or defects and enabling automated deployments, freeing up time spent manually building, testing, or releasing software.

CI/CD introduces automated processes to integrate code and test in a continuous manner with delivery and deployment, replacing error-prone manual processes. CI/CD is supported by teams working together in an agile way, either with DevOps or SRE practices.

5. Version control

Version control or revision control systems help software developers keep track of changes to application code and manage the development of a single program by more than one person. Version control systems such as Git have the ability to create branches, where a developer can make a copy of an existing project and modify one or more files.

6. NoSQL databases

NoSQL databases are a class of database management systems (DBMSs) that do not rely on the traditional relational database management system (RDBMS) structure. NoSQL databases are purpose-built for specific data models, have flexible schemas for building modern applications, and are widely recognized for their ease of development and performance at scale. These databases use various data models for accessing and managing data, which makes them optimized specifically for applications that require large data volume, low latency, and flexible data models.

Interested in remote Site Reliability jobs?

Become a Turing developer!

Apply now

How to get remote Site Reliability engineer jobs?

Developers are a lot like athletes. In order to excel at their craft, they have to practice effectively and consistently. They also need to work hard enough so that their skills grow gradually over time. In that regard, there are two major factors that developers must focus on in order for that progress to happen: the support of someone who is more experienced and effective in practice techniques while you're practicing. As a developer, it's vital for you to know how much to practice - so make sure there is someone on hand who will help you out and keep an eye out for any signs of burnout!

Turing offers the best remote site reliability engineer jobs that suit your career trajectories as a site reliability engineer. Grow rapidly by working on challenging technical and business problems on the latest technologies. Join a network of the world's best developers & get full-time, long-term remote site reliability engineer jobs with better compensation and career growth.

Why become a Site Reliability engineer at Turing?

Elite US jobs
Career growth
Exclusive developer community
Once you join Turing, you’ll never have to apply for another job.
Work from the comfort of your home
Great compensation

How much does Turing pay their Site Reliability engineers?

At Turing, every site reliability engineer is allowed to set their rate. However, Turing will recommend a salary at which we know we can find a fruitful and long-term opportunity for you. Our recommendations are based on our assessment of market conditions and the demand that we see from our customers.

Frequently Asked Questions

Turing is an AGI infrastructure company specializing in post-training large language models (LLMs) to enhance advanced reasoning, problem-solving, and cognitive tasks. Founded in 2018, Turing leverages the expertise of its globally distributed technical, business, and research experts to help Fortune 500 companies deploy customized AI solutions that transform operations and accelerate growth. As a leader in the AGI ecosystem, Turing partners with top AI labs and enterprises to deliver cutting-edge innovations in generative AI, making it a critical player in shaping the future of artificial intelligence.

After uploading your resume, you will have to go through the three tests -- seniority assessment, tech stack test, and live coding challenge. Once you clear these tests, you are eligible to apply to a wide range of jobs available based on your skills.

No, you don't need to pay any taxes in the U.S. However, you might need to pay taxes according to your country’s tax laws. Also, your bank might charge you a small amount as a transaction fee.

We, at Turing, hire remote developers for over 100 skills like React/Node, Python, Angular, Swift, React Native, Android, Java, Rails, Golang, PHP, Vue, among several others. We also hire engineers based on tech roles and seniority.

Communication is crucial for success while working with American clients. We prefer candidates with a B1 level of English i.e. those who have the necessary fluency to communicate without effort with our clients and native speakers.

Currently, we have openings only for the developers because of the volume of job demands from our clients. But in the future, we might expand to other roles too. Do check out our careers page periodically to see if we could offer a position that suits your skills and experience.

Our unique differentiation lies in the combination of our core business model and values. To advance AGI, Turing offers temporary contract opportunities. Most AI Consultant contracts last up to 3 months, with the possibility of monthly extensions—subject to your interest, availability, and client demand—up to a maximum of 10 continuous months. For our Turing Intelligence business, we provide full-time, long-term project engagements.

No, the service is absolutely free for software developers who sign up.

Ideally, a remote developer needs to have at least 3 years of relevant experience to get hired by Turing, but at the same time, we don't say no to exceptional developers. Take our test to find out if we could offer something exciting for you.

View more FAQs

Latest posts from Turing

Software-developer-jobs-in-Silicon-Valley-tech-companies

Looking for Software Developer Jobs? Learn How to Write a Clean Code First

Are you a software developer looking for remote jobs in Silicon Valley tech companies? If yes, these clean code t...

Read more

Turing.com Review by Nigeria’s Joy: Flexibility in Work Allows Me to Enjoy Life More

Flexibility in work at Turing allows me to enjoy life more, says Nigeria’s Joy in her Turing.com review...

Read more
Software-Development-Life-Cycle-scaled

The Nine Steps of Software Product Development Life Cycle

A product development process depends on the nature of the business. But these steps can turn your ordinary softw...

Read more

Ten Tips to Crack a Software Developer Job Interview

Cracking a software developer job interview is no cakewalk. Here are a few tips to help level up your...

Read more

Turing Blog: Articles, Insights, Company News and Updates

Explore insights on AI and AGI at Turing's blog. Get expert insights on leveraging AI-powered solutions to drive ...

Read more

Leadership

In a nutshell, Turing aims to make the world flat for opportunity. Turing is the brainchild of serial A.I. entrepreneurs Jonathan and Vijay, whose previous successfully-acquired AI firm was powered by exceptional remote talent. Also part of Turing’s band of innovators are high-profile investors, such as Facebook's first CTO (Adam D'Angelo), executives from Google, Amazon, Twitter, and Foundation Capital.

Equal Opportunity Policy

Turing is an equal opportunity employer. Turing prohibits discrimination and harassment of any type and affords equal employment opportunities to employees and applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, age, disability status, protected veteran status, or any other characteristic protected by law.

Explore remote developer jobs

briefcase
Principal AI Engineer - US NYC

Principal AI Engineer

Location: US(NYC)- 3(WFO)

Employment Type: Full Time (Overlapping EST)

Experience Level: Staff/Principal (8–14 years)


About the Role

Turing is hiring a Staff/Principal AI Engineer to lead enterprise-scale agentic AI implementations for Fortune 500 clients. This is a hands-on engineering role focused on designing and shipping autonomous, tool-calling AI systems — agents that reason over enterprise context, invoke real systems through secure interfaces, and operate reliably at scale under strict latency, cost, and governance constraints.

You will own these systems end to end: the data pipelines feeding them, the backend services around them, the agent orchestration layer, the evaluation harness that keeps them honest, and the cloud infrastructure they run on. We are looking for engineers with genuine software engineering and data science depth who have taken agentic systems all the way to production.

What We're Looking For

Engineering foundation

  • 8–14 years of software engineering experience, with strong hands-on large-scale Python
  • Working depth in at least one systems or backend language — Go, Rust, Java, or C/C++ — and the judgment to know when to reach for it
  • Strong data structures and algorithms.
  • Strong understanding of APIs, microservices, and system design
  • Hands-on experience building and operating data pipelines and production-grade distributed systems.

Agentic AI and LLMs

  • 2+ years of hands-on LLM engineering, with at least couple agentic system you designed and took to production
  • Production experience with agent frameworks — LangGraph, Google ADK, CrewAI, Claude Agent SDK, or equivalent — and the fluency to move between them as the ecosystem evolves
  • Experience building MCP (Model Context Protocol) servers and tool-calling interfaces
  • RAG from first principles: chunking strategy, embeddings, vector and hybrid retrieval, reranking, and response validation
  • Strong experience with vector databases (Milvus, Pinecone, Weaviate, FAISS, etc. or cloud equivalents)
  • Design of guardrails and reliability patterns — validators, policy checks, self-correction loops, deterministic fallbacks, circuit breakers, and rollback paths

Optimization

  • Deep familiarity with token optimization and context-window management — context shaping, pruning, and compaction
  • Latency and cost optimization through caching, model routing, batching, streaming, and parallel tool calls
  • Performance testing and tuning systems against defined SLOs

Evaluation

  • Experience building evaluation frameworks for LLM systems — offline eval sets, continuous online evaluation, and regression detection
  • Instrumentation and traceability suitable for regulated enterprise environments using tools like LangSmith, Langfuse, etc.

Cloud

  • Hands-on AWS: containerized services (ECS/EKS), serverless (Lambda), data services (S3, DynamoDB, Redshift) and orchestration (Step Functions); Azure or GCP equivalents also valued
  • Familiarity with CI/CD pipelines and DevOps practices
  • Infrastructure as code with Terraform or CloudFormation, and mature CI/CD practice

Working traits

  • Strong analytical problem-solving with a bias to ownership and urgency
  • Clear cross-team communication, working directly with client stakeholders to translate business problems into technical roadmaps
  • Able to work productively in ambiguity from system-level documentation and ramp quickly in unfamiliar codebases

Good to Have

  • Experience with managed AI platforms — Amazon Bedrock, Vertex AI, Azure AI — paired with fluency in the underlying fundamentals

Roles & Responsibilities

  • Design and build agentic systems: Lead the architecture and implementation of tool-calling agents that combine retrieval, structured reasoning, and secure action execution with least-privilege access.
  • Productionize LLM applications: Build retrieval pipelines, prompt synthesis, response validation, and self-correction loops, backed by rigorous evaluation.
  • Own the full stack: Deliver the data pipelines, backend services, distributed compute, and orchestration layer that agentic systems depend on — not only the model invocation.
  • Engineer for reliability and governance: Build validator models, adversarial test suites, and policy checks; enforce deterministic fallbacks and rollback strategies; instrument continuous evaluation.
  • Optimize for cost and latency: Drive measurable improvements in token efficiency, response time, and unit economics against defined SLOs.
  • Codebase ownership: Build, maintain, and review high-quality Python and SQL, with an emphasis on reusable components, scalability, and performance.
  • Cloud integration: Deploy AI applications on AWS, Azure, or GCP with optimized resource usage and robust CI/CD.
  • Cross-functional collaboration: Partner with product owners, data scientists, and business SMEs to define requirements and deliver impactful AI products.
  • Mentoring and technical leadership: Set engineering standards and share knowledge across the team, raising the bar on AI and software engineering practice.
Finance
10K+ employees
PythonGoRust+ 4
briefcase
Lead Edge AI & Computer Vision Engineer

Lead Edge AI & Computer Vision Engineer

  • Location: USA [with regular on-site travel required to client site]
  • Target Start Date: 1 Sep
  • Time Commitment: 8 weeks full-time
  • Experience Level: 12–15+ Years

Core Objective: Architect, benchmark, and optimize an end-to-end computer vision and low-latency execution pipeline—combining spatial geometric math with low-level C++/CUDA acceleration directly on NVIDIA Jetson embedded edge hardware (Orin NX / AGX Orin) to achieve sub-10 ms processing latency.

Key Responsibilities

  1. Model Selection & Spatial Math: Evaluate real-time detection topologies (e.g. YOLO) at 100+ FPS and build 2D perspective homography unwarping, lens undistortion, and spatial algorithms to convert pixel coordinates into physical millimeter units within tight error bounds (millimeter level accuracy).
  2. TensorRT INT8 Acceleration: Execute Post-Training Quantization to compile PyTorch/ONNX models into high-throughput TensorRT INT8/FP16 engines on Jetson Orin hardware without accuracy degradation.
  3. Zero-Copy Memory Architecture:  Engineer zero-copy memory pipelines using NVIDIA Memory Management and DMA transfers to eliminate bottlenecks.
  4. Compiled C++ Execution & I/O: Build compiled C++17 execution frameworks with multi-threaded lock-free ring buffers and non-blocking asynchronous socket communication.
  5. Nsight Profiling & Roadmapping: Instrument stage-by-stage pipeline latency using NVIDIA Nsight Systems/NVTX markers to bound tail latency, author feasibility reports, and design technical roadmaps.

Key Qualifications & Experience

  1. Full-Stack Edge AI Experience: 12–15+ years of experience bridging real-time computer vision algorithm design, spatial geometry, and low-level C++/CUDA execution on embedded edge hardware.
  2. NVIDIA Jetson Ecosystem: Deep expertise in Jetson embedded platforms (Orin NX, AGX Orin, L4T, power profiles, core pinning) and a proven track record compiling/tuning TensorRT FP16/INT8 engines via PTQ/QAT.
  3. Optical & Spatial Geometry: Expertise in 2D/3D camera calibration, perspective homography transformations, lens distortion modeling, and millimeter sizing math in industrial settings.
  4. Low-Latency Systems Engineering: Expertise in zero-copy shared memory, DMA frame buffers, lock-free queues, custom CUDA plugins, and microsecond profiling via Nsight Systems and NVTX markers.
  5. Tooling & Industrial I/O: Proficiency in C++17, Python, PyTorch, OpenCV, CUDA, non-blocking asynchronous sockets (UDP, PLC integration), and building automated dataset benchmarking harnesses.
Manufacturing
10K+ employees
C++CUDANVIDIA Omniverse+ 5
sample card

Apply for the best jobs

View more openings
Turing books $87M at a $1.1B valuation to help source, hire and manage engineers remotely
Turing named one of America's Best Startup Employers for 2022 by Forbes
Ranked no. 1 in The Information’s "50 Most Promising Startups of 2021" in the B2B category
Turing named to Fast Company's World's Most Innovative Companies 2021 for placing remote devs at top firms via AI-powered vetting
Turing helps entrepreneurs tap into the global talent pool to hire elite, pre-vetted remote engineers at the push of a button

Work with the world's top companies

Create your profile, pass Turing Tests and get job offers as early as 2 weeks.