/#jobs
#Platform#About Us#Employers#AI Jobs
Get Started
Sydney
1 day ago
AI Engineer - Inference logo

AI Engineer - Inference

Firmus Technologies

Sydney
1 day ago
Apply

AI Engineer - Inference

Firmus Technologies is seeking a Senior AI Engineer (Inferencing) to build and improve the AI & Applications team's inference capability, making models available as reliable, secure, scalable, and high-performance endpoints. The role involves establishing the engineering foundation for self-hosted model serving, optimizing performance, and contributing to the Model-to-Grid product and agentic applications roadmap. Requires 5+ years of software engineering experience, including 3+ years in AI inference or related fields, and hands-on experience with modern inference frameworks.

Core AIHybridFull-timeSeniorTensorRT-LLMTensorRT

AI Engineer - Inference

Firmus Technologies is seeking a Senior AI Engineer (Inferencing) to build and improve the AI & Applications team's inference capability, making models available as reliable, secure, scalable, and high-performance endpoints. The role involves establishing the engineering foundation for self-hosted model serving, optimizing performance, and contributing to the Model-to-Grid product and agentic applications roadmap. Requires 5+ years of software engineering experience, including 3+ years in AI inference or related fields, and hands-on experience with modern inference frameworks.

Apply
Core AIHybridFull-timeSeniorTensorRT-LLM

Salary

Not specified

Work Location

Sydney, New South Wales, Australia, AU

Work Model

Hybrid

Experience Required

5 years

Employment Type

Full-time

Experience Level

Senior

Core Qualifications

Technical (Must-have)
TensorRT-LLMTensorRTSGLangvLLMTriton Inference ServerNVIDIA DynamoNVIDIA NIMCUDAcuDNNNCCLPythonC++GoKubernetesCI/CD
Soft Skills
CollaborationCommunicationProblem SolvingCuriosityDrive

Key Responsibilities

  • •Build, operate, and continuously improve self-hosted AI inference services for internal applications, customer-facing products, and future Inference-as-a-service offerings.
  • •Define and implement standard model-onboarding workflows covering model intake, compatibility validation, packaging, runtime selection, optimization, deployment, endpoint registration, testing, release, and lifecycle management.
  • •Provision and manage secure, scalable inference endpoints for common AI application patterns, including interactive generation, RAG, embeddings, reranking, batch processing, multimodal use cases, tool calling, and agentic workflows.
  • •Develop reusable deployment templates, APIs, SDKs, configuration standards, and self-service workflows for users to request, configure, access, monitor, update, and retire model endpoints.
  • •Work with leading inference frameworks and toolkits, such as TensorRT-LLM, TensorRT, SGLang, vLLM, Triton Inference Server, NVIDIA Dynamo, NVIDIA NIM, CUDA, cuDNN, NCCL, and related serving, profiling, and observability tools.
  • •Optimize model-serving performance using appropriate techniques, including quantization, compilation, batching, continuous batching, request routing, KV-cache management, prefix caching, speculative decoding, load balancing, model routing, memory optimization, and distributed parallelism.
  • •Build and validate reusable inference recipes that specify compatible model versions, framework and runtime versions, precision formats, GPU configurations, topology requirements, scaling approaches, scheduler profiles, benchmark results, and expected performance envelopes.
  • •Use quantization and optimization approaches such as NVFP4, FP8, INT8, TensorRT compilation, kernel optimization, efficient attention mechanisms, and memory-management techniques while maintaining agreed model-quality targets.
  • •Design distributed inference configurations for large models, including tensor, pipeline, expert, context, and data parallelism where appropriate.
  • •Work with the Kubernetes and proprietary scheduler team to define endpoint resource profiles, placement requirements, topology preferences, priority classes, quota models, autoscaling rules, capacity reservations, and workload-management policies.
  • •Contribute inference workload characteristics, benchmarks, and performance profiles to the Model-to-Grid product so that endpoint placement, scheduling, capacity planning, and AI-factory operations can make more informed decisions.
  • •Build benchmarking and qualification workflows using controlled experiments, reproducible baselines, load tests, latency tests, throughput tests, concurrency tests, scaling tests, performance profiling, regression testing, and internal or industry-standard benchmark methodologies where relevant.
  • •Measure and improve key inference indicators, including time-to-first-token, inter-token latency, tokens per second, requests per second, end-to-end latency, concurrency, GPU utilization, memory efficiency, cache hit rate, scaling efficiency, power efficiency, and cost efficiency.
  • •Establish automated performance-regression testing and release qualification for model versions, runtime and toolkit upgrades, CUDA and driver changes, Kubernetes releases, scheduler changes, networking and storage changes, and new GPU platforms.
  • •Build operational observability for inference services, including endpoint availability, request volume, latency, queueing, errors, GPU utilization, GPU memory use, cache behavior, capacity, cost, power, and service-level objectives.
  • •Partner with the agentic applications team to provide fit-for-purpose self-hosted endpoints for agent planning, retrieval, tool use, summarization, diagnosis, recommendation, optimization, and AI-factory operations.
  • •Expose governed inference, benchmark, recipe, performance, and capacity information to agentic systems, allowing them to recommend suitable models, identify degradation, diagnose bottlenecks, plan optimization experiments, and validate results.
  • •Work with Product, UX, DevOps, Platform, Infrastructure, Security, and Global Operations teams to ensure that inference provisioning, model selection, endpoint configuration, performance visibility, quota management, and troubleshooting are clear, secure, and operationally supportable.
AI EngineerInferenceGPUModel ServingLLMKubernetesNVIDIASeniorHybridFull-time
/#jobs

Your gateway to a successful career. Show your growth. Be ready for your next step. Capture and seize the best opportunities.

  • Data
  • FAQ
  • Articles
  • AI Jobs
  • Platform
  • Employers
  • About Us
  • Legal
© 2026/#jobsAll rights reserved.

For queries/support, email jobs.support@slashhash.ai