BE-2026Contact

Berk ErdoganIndependent developer for quantized LLMs, local inference, and native apps

Based in Istanbul. Open-source model builds on Hugging Face, a self-hosted one-million-token inference server, and apps on the App Store.

58,844all-time downloads of 5 open-source models

Select a block on the die to zoom in. Esc zooms out.Tap a block on the die to zoom in.

Services

I am an independent developer in Istanbul. I quantize open-weight language models for NVIDIA Blackwell GPUs and publish them on Hugging Face, where they have more than 58,000 downloads. I build and run my own inference server, and I ship native apps to the App Store.

  • P0 · BE-2026-APP

    iOS and Android apps

    Native SwiftUI and Kotlin apps, from the first screen to App Store release: StoreKit 2 subscriptions, App Store Connect, and Search Ads.

  • P1 · BE-2026-LLM

    LLM quantization

    NVFP4, FP8, and INT8 builds of open-weight models, with calibration-set design and KL / ΔNLL evaluation against the BF16 source.

  • P2 · BE-2026-LLM

    Private LLM inference

    Local GPU servers with vLLM: multi-GPU serving, long context, FP8 KV cache, and API endpoints for your own tools.

  • P3 · BE-2026-WEB

    Websites and business software

    TypeScript sites on Cloudflare, and desktop business apps in SwiftUI or Electron.

Proof

  • 5 open-source quantized LLMs, 58,844 all-time downloads
  • 1,000,000-token context validated on 2× RTX 5090: 4/4 needle retrievals at 989k tokens
  • 2 apps released on the App Store, with subscriptions and paid acquisition
  • Business software in daily use at a 5-employee firm

Start

Tell me what you want to build. I reply by email.

Request a quote

Models

Published quantizationsDownloads: all-time, read live from the Hugging Face Hub
Published quantized models
ModelBase modelFormatToolchainReleasedDownloads
Qwen3.6-27B-NVFP4Qwen3.6-27B (vision-language) · NVIDIA ModelOpt · 2026-05-02Qwen3.6-27B (vision-language)NVFP4 W4A4NVIDIA ModelOpt2026-05-0249,909
Qwen3.8-27B-NVFP4-1MQwen3.8-27B · llm-compressor, GPTQ · 2026-08-20Qwen3.8-27BNVFP4 + FP8, FP8 KV cachellm-compressor, GPTQ2026-08-20562
Qwen3.8-27B INT8 AutoRound (community variant)Qwen3.8-27B community variant · AutoRound (SignRound) · 2026-08-23Qwen3.8-27B community variantINT8 W8A16, group 128AutoRound (SignRound)2026-08-231,310
gemma-4-12B-it-NVFP4gemma-4-12B-it (multimodal) · NVIDIA ModelOpt · 2026-06-04gemma-4-12B-it (multimodal)NVFP4NVIDIA ModelOpt2026-06-043,978
Qwen3.5-27B-NVFP4Qwen3.5-27B (multimodal) · NVIDIA ModelOpt · 2026-03-10Qwen3.5-27B (multimodal)NVFP4NVIDIA ModelOpt2026-03-103,085
Total58,844
All-time downloads vs release date
1001k10k100kMarAprMayJunJulAugSepDownloads (log scale)Release date, 2026Qwen3.6-27B-NVFP449,909 · 2026-05-02Qwen3.8-27B-NVFP4-1M562 · 2026-08-20Qwen3.8-27B INT8 AutoRound (community variant)1,310 · 2026-08-23gemma-4-12B-it-NVFP43,978 · 2026-06-04Qwen3.5-27B-NVFP43,085 · 2026-03-10

The plot is for wider screens. The same values are in the table above.

Fine-tuned modelsReserved

GPU 6–10 · reserved for fine-tuned models

  • GPU 6Reserved for a fine-tune
  • GPU 7Reserved for a fine-tune
  • GPU 8Reserved for a fine-tune
  • GPU 9Reserved for a fine-tune
  • GPU 10Reserved for a fine-tune

Fine-tuned models are in preparation. They will appear here and on huggingface.co/berkerdooo when the weights are published.

Inference server
Inference server ratings
ParameterValue
Context length, validated (1)1,000,000 tokens
GPU memory, 2× RTX 509064 GB
System memory, DDR5128 GB
Tensor parallel degree2 GPUs
InterconnectPCIe

(1) Qwen3.8-27B-NVFP4-1M, end to end: 4 of 4 needle-in-a-haystack retrievals at 989k prompt tokens.

Serving configuration
  • vLLM, tensor parallel over two GPUs
  • FP8 KV cache with calibrated static scales
  • FlashInfer attention and prefix caching
  • MTP speculative decoding
  • Native Anthropic Messages API endpoint, used as the backend for Claude Code and opencode
  • llama.cpp and GGUF every day on Apple Silicon and CUDA
Quality characteristics
Quality characteristics of the published models
ParameterTest conditionsTypical
Top-1 agreement vs BF16INT8 AutoRound W8A16, teacher-forcedINT8 AutoRound W8A16, teacher-forced97.3 %
Needle retrievalNVFP4-1M, 989k prompt tokens, 2× RTX 5090NVFP4-1M, 989k prompt tokens, 2× RTX 50904/4 hits
Calibration mixChat, prose, code and mathChat, prose, code and math512 × 4,096 tokens
Evaluation streamTop-24 KL divergence and ΔNLL vs BF16 and FP8 referencesTop-24 KL divergence and ΔNLL vs BF16 and FP8 references~1M tokens
Agentic benchmarkDeepSWE, long-horizon tasks in Docker sandboxesDeepSWE, long-horizon tasks in Docker sandboxes113 tasks

Apps

  • KalemPRODUCTION

    A Qur’an app that answers questions only from the Qur’an and its commentary, and shows the verses behind every answer. Prayer times, Qibla, hadith library, memorization. English, Turkish, and French.

    SwiftUI · StoreKit 2 subscriptions · on-device tafsir · Apple Search Ads

    Platform
    iOS
    Released
    Jan 2026
    View on the App Store
  • Sivas TractorPRODUCTION

    A marketplace for new and used tractors and farm equipment. A native SwiftUI rewrite of the original Flutter app, with a Kotlin/Compose Android port and a web storefront.

    SwiftUI · Kotlin/Compose · Firebase Auth + Firestore rules · Cloudinary

    Platform
    iOS · Android · Web
    Released
    Mar 2026
    View on the App Store
  • ReelmIN DEVELOPMENT

    An AI video-generation app with a credit economy.

    SwiftUI · Swift 6 concurrency · Firebase · StoreKit 2 credits · Fal AI

    Platform
    iOS
    Released
    In development

Skills

16-pin package, top view
Skills and tools
No.NameTypeDescription
1PYLANGPython: quantization, calibration, and evaluation pipelines
2SWIFTLANGSwift and SwiftUI: iOS and macOS apps, Swift 6 concurrency
3KTLANGKotlin and Jetpack Compose: Android apps
4TSLANGTypeScript and JavaScript: web, Electron, Workers
5SQLLANGSQL: SQLite, Postgres
6TORCHMLPyTorch and Transformers
7QUANTMLllm-compressor, NVIDIA ModelOpt, AutoRound: NVFP4, FP8, INT8, GPTQ, AWQ
8EVALMLKL divergence and ΔNLL evaluation, calibration-dataset design
9VLLMSERVEvLLM multi-GPU serving, FP8 KV cache, speculative decoding
10GGUFSERVEllama.cpp and GGUF on Apple Silicon and CUDA
11CUDASERVEMulti-GPU CUDA inference on Linux
12DOCKEROPSDocker, Linux server administration, Git, shell
13BAASPLATFORMFirebase and Supabase
14CFPLATFORMCloudflare Workers and Pages
15STOREPLATFORMStoreKit 2 and App Store Connect operations
163DTOOLBlender, 3ds Max, ZBrush, AutoCAD, Substance Painter

About

I studied Industrial Engineering at Istanbul Technical University. Before that, I studied English in London and in Toronto. Industrial engineering still shapes how I work: I measure first, find the bottleneck, and ship on schedule.

I taught myself software by building things I needed. Today my work has two tracks. The first is making large open-weight models run well on local GPUs. The second is shipping native apps to the App Store and building software for small businesses.

  • EducationBSc Industrial Engineering, Istanbul Technical University, 2019–2024
  • Language studyEmbassy English: Toronto, 2018 · London, 2012
  • 2023–2025Industrial engineer, Erdogan Architecture: interior projects for clients in Switzerland and France, 3D visualization, site management
  • InternshipsDeha Advertisement (marketing) · Ferre / Femas Metal (production)
  • Outside work3D art · DJ and music production
  • LanguagesTurkish (native) · English (C2) · German (A1)

Client work

Custom designsClient names withheld
Client work
DesignDescription
Accounting software for macOSSwiftUI and SwiftData, a 261-test suite, and an Excel reporting engine. In daily use at a 5-employee firm.
Offline cash-bookElectron desktop app with an encrypted SQLite store. In daily use at the same firm.
Websites and web storefrontsTypeScript sites on Cloudflare Workers and Pages, with Firebase or Supabase back ends.

Contact

What you can order
PartDescription
BE-2026-APPCustom iOS or Android app, from design to App Store release
BE-2026-LLMModel quantization, evaluation, or private local inference setup
BE-2026-WEBWebsite or business software
BE-2026-HIREFull-time or contract role

Write to [email protected], or use the form.