Prajjwal C.

Prajjwal C.

George Town, Cayman Islands
5K followers 500+ connections

About

Backend engineer building crypto-native, money-critical infrastructure. I work where financial correctness meets scale: on-chain/off-chain consistency, settlement accuracy, idempotency, and the boring-but-load-bearing systems that move real value safely.
Currently at ether.fi, building backend systems for a multi-billion-dollar liquid restaking protocol.
Previously:
Stader Labs - founding engineer; helped scale liquid-staking infrastructure past $500M+ TVL across multiple chains.
Visa - Senior SDE on payments-scale distributed systems, where reliability and correctness under load were non-negotiable.
Across all of it the throughline is the same: high-throughput systems where getting the money math wrong is not an option. Smart-contract integration, exchange and settlement flows, financial invariants, failure modes, and observability for systems handling real user funds.
Stack: Go, Rust, Solidity, Java, Kafka, distributed systems.
On the side I go deep on low-level performance and correctness in the LLM inference ecosystem (vLLM, NVIDIA KVPress, constrained decoding), contributing open source. Same instincts as distributed systems: throughput, memory, latency, correctness.
Competitive programming: Codeforces Expert (~8000 problems solved).

Activity

5K followers

See all activities

Experience

  • ether.fi

    George Town, Cayman Islands

  • -

    Bengaluru, Karnataka, India

  • -

    Bengaluru, Karnataka, India

  • -

  • -

    Bangalore Urban, Karnataka, India

  • -

    Singapore

  • -

    Remote

  • -

    New Delhi, Delhi, India

  • -

    New Delhi, Delhi, India

  • -

    New Delhi, Delhi, India

  • -

    Delhi, India

  • -

    Pune Division, Maharashtra, India

  • -

    West Bengal, India

  • -

    Manipal

Education

Licenses & Certifications

Volunteer Experience

  • Open Source Developer

    qxresearch

    - Present 5 years 5 months

    Science and Technology

  • Member

    NVIDIA GameWorks

    Science and Technology

  • Member

    DWYL

    - Present 5 years 5 months

    Science and Technology

  • Open Source Contributer

    EddieHub

    - Present 5 years 9 months

    Science and Technology

Publications

  • Perplexity Holds, Programs Break: Executable Correctness as a Blind Spot of KV-Cache Compression

    Zenodo

    KV-cache compression methods are evaluated almost exclusively with metrics that score token or string overlap, retrieval hits, or an extracted final answer. None of the widely used long-context suites check whether generated output actually runs, or whether a generated tool call is well-formed. We argue this is a systematic blind spot: lexical and retrieval metrics can stay flat while a compressed cache quietly breaks structured generation, because a single dropped or garbled token fails a unit…

    KV-cache compression methods are evaluated almost exclusively with metrics that score token or string overlap, retrieval hits, or an extracted final answer. None of the widely used long-context suites check whether generated output actually runs, or whether a generated tool call is well-formed. We argue this is a systematic blind spot: lexical and retrieval metrics can stay flat while a compressed cache quietly breaks structured generation, because a single dropped or garbled token fails a unit test or a schema check but barely moves a substring score. We introduce kv-exec-bench, a small open benchmark that measures the two missing quantities, code unit-test pass@1 and tool-call JSON-Schema validity, under KV compression. It is built on top of NVIDIA's kvpress as a dependency, so any press works without modification. In a deliberately small, CPU-only study on a 0.5B-parameter model, we observe the predicted divergence on tool calling: for the best-behaved presses, the rate of parseable tool calls stays at 100 percent across compression ratios while the rate of schema-valid calls falls by half or more. A parse-only or string-only metric would report this regime as nearly lossless. We release the benchmark and all code under Apache-2.0, and frame this as a measurement-gap report plus a reusable benchmark rather than a large-scale empirical study.

    See publication
  • StragglerPolicy: Straggler-Aware Elastic Membership for Decentralized Training

    Zenodo

    Decentralized training methods such as DiLoCo make low-communication language-model training over commodity, geographically distributed hardware practical, and production stacks (e.g. Prime Intellect's prime-diloco/PCCL) already tolerate dead nodes through heartbeat eviction and elastic join/leave. They do not, however, handle the slow-but-alive straggler: a single node running at a fraction of peer throughput stalls every synchronous outer step, because the barrier waits for everyone. We…

    Decentralized training methods such as DiLoCo make low-communication language-model training over commodity, geographically distributed hardware practical, and production stacks (e.g. Prime Intellect's prime-diloco/PCCL) already tolerate dead nodes through heartbeat eviction and elastic join/leave. They do not, however, handle the slow-but-alive straggler: a single node running at a fraction of peer throughput stalls every synchronous outer step, because the barrier waits for everyone. We present StragglerPolicy, a membership policy that adds an adaptive per-round soft deadline (median + k*MAD over a rolling arrival-offset history), a partial-participation quorum with sidelining, and a graduated slow-node response (transient slow -> rejoin; persistently slow -> evict). On a persistent-straggler scenario (four workers, one 10x slow), StragglerPolicy is 4.59x faster than a faithful PRIME/PCCL-style baseline in a deterministic discrete-event simulator, raising worker utilization from 0.62 to 0.89. To establish that this is not a simulation artifact, we re-run both policies unchanged on a real torch.distributed/gloo DiLoCo loop; the straggler win reproduces directionally (1.32x faster on real code, slow rank sidelined four times then evicted), and a simulator calibrated to the real run's constants predicts the same ordering. We are explicit about the gap between the simulator's relative speedup and the absolute speedup observed on real hardware, and frame the simulator as a tool for relative comparison of membership policies rather than an absolute-throughput predictor. The simulator, policy, and validation harness are open source.

    See publication

Courses

  • Algorithm Design And Analysis

    -

  • Compiler Design

    -

  • Computer Architechture

    -

  • Computer Networks

    -

  • Data Structures

    -

  • Database Management Systems

    -

  • Machine Learning

    -

  • Object Oriented Software Engineering

    -

  • Operating Systems

    -

Projects

  • 6502 Processor emulator

    -

    • Emulated a 6502-processor architecture with all defined pins
    definition and processor addressing modes.
    • Defined all of the 56 operation codes, Bus connections
    emulated with a header file written in C++.

    See project
  • Bellman Ford Visualization

    -

    About
    Simulation of Bellman Ford Algorithm using GLUT openGL.Dijkstra’s algorithm is a Greedy algorithm and time complexity is O(VLogV) . Implemented using freeglut libraries with cost matrix as input. can be used on negative weighted graphs too.

    See project
  • Dark Theme Materials Design Android Apps ( Memo+ and Calculator)

    -

    https://github.com/pjdurden/memo-
    https://github.com/pjdurden/Calculate-flutter
    A materials design dark theme calculator App built as a project for 30daysofflutter (Google) using Flutter SDK.
    memo adding and saving app for Android written in Kotlin

    See project
  • restake-xray

    -

    Open-source Go engine that maps restaking / liquid-restaking-token (LRT) exposure across protocols. Author and maintainer.

  • Tic-Tac-Toe using Socket Programming

    -

    A Tic-Tac-Toe game to play against computer using Min-Max
    Algorithm further modified with Alpha Beta Pruning an
    Artificial Intelligence game search algorithm to predict the
    best possible move

    See project
  • Video/Image Analysis Using Machine Learning

    -

    Implementation of YOLO algorithm to do Video and image
    analysis build on Google colab notebook.Implemented
    Convolutional Neural Network written in C and CUDA.

    See project
  • Vidzz Android Application

    -

    An application to record, share, and browse short videos on
    an area-based user feed utilizing Firebase Real-time Database
    and Storage for fast, scalable streaming of videos(implemented
    cache management).

    See project

Honors & Awards

  • CodeChef Best Finishes 2020

    -

    Highest Rating - 1890
    February Challenge 2021 Div 2: Rank - 52 ( College Rank 2)
    January Challenge 2021 Div 3: Rank - 439 (College Rank 12)
    CodeChef push_back(2): Rank - 382 (College Rank 6)
    January Lunchtime 2021 Div 2: Rank - 636

  • First Runnerup at Codesprint

    -

    Secured rank 2 at Codesprint organised by Developer Student Club of Bharati Vidyapeeth College of Engineering

View Prajjwal’s full profile

  • See who you know in common
  • Get introduced
  • Contact Prajjwal directly
Join to view full profile

Other similar profiles

Explore collaborative articles

We’re unlocking community knowledge in a new way. Experts add insights directly into each article, started with the help of AI.

Explore More

Add new skills with these courses