Services/AI Development
AI Development in San Diego: agents, RAG, and on-device models that ship
DevX Group LLC builds AI features that ship: agents that run real workflows, RAG systems that answer from your own documents, and on-device models that keep data on the machine. Based in San Diego and led by founder Max Sheikhizadeh, priced at $95 to $175 per hour, with most first releases delivered in 4 to 12 weeks.
What we build
Most AI projects fail at the point where a demo has to become a feature people rely on. We start from the task, not the model: which decision or document or hour of manual work is worth automating, and how will you know it worked.
- AI agents and workflow automation: an agent that reads the inbox, triages the ticket, drafts the reply, files the record, and asks a human only when confidence is low.
- RAG systems: a knowledge assistant that answers from your contracts, manuals or support history, cites the source, and says "I do not know" instead of inventing.
- On-device and local AI: speech, vision and language models that run on the Mac or the phone, so the data never leaves it. Our own Parlin is built this way.
- AI inside an existing product: a smart feature added to the app you already have, with the evaluation set, cost controls and fallbacks that keep it dependable.
How we keep it honest
Every engagement begins with an evaluation set: 30 to 100 real examples of the task, with the answer a good employee would give. We measure the first prototype against it in week two, and every change after that has to move the score. No vibes, no "it seems to work."
Cost is designed in, not discovered on the invoice. We size the model to the task, cache what repeats, and put a spend ceiling on every agent before it touches production.
Who this is for
Founders with a product that needs one AI feature done well, operations leads drowning in repetitive work, and teams that tried a chatbot builder and hit its ceiling. If you are in San Diego we can meet in person; most of our clients work with us remotely across US time zones.
What we build with
Named tools, and the reason each one is on the list.
Claude and GPT models
Frontier models through their APIs, picked per task and swappable.
Vercel AI SDK
One provider abstraction, so a model change is a config change.
Supabase and pgvector
Retrieval, auth and data in one Postgres, no separate vector database to run.
Next.js and TypeScript
The web layer every agent and assistant ships inside.
Ollama, Core ML and on-device models
Local inference for privacy-first products and offline use.
Evaluation harnesses
Scored test sets run on every change, the same way unit tests guard code.
How an engagement runs
Bi-weekly demos throughout. Code, configs and runbooks are yours at acceptance.
- 01Week 1
Discovery and evaluation set
We pick the one task worth automating, collect real examples, and agree what "correct" means before any model is called.
- 02Weeks 2 to 4
Prototype on real data
A working prototype scored against the evaluation set, shown to you in a bi-weekly demo. Model, prompt and retrieval choices are made here, with the numbers.
- 03Weeks 4 to 8
Hardening
Guardrails, fallbacks, cost ceilings, logging and the human hand-off path. This is the part chatbot builders skip and the reason their demos never ship.
- 04Weeks 8 to 12
Launch and 30-day stabilization
Production release with monitoring, then 30 days of on-call fixes and tuning while real usage arrives. Code, configs and runbooks are yours at acceptance.
Proof, not promises
Shipped work that shows this service in practice.
Parlin: Private Mac Transcription and Dictation
Parlin runs transcription, dictation and translation fully on the Mac, no audio ever uploaded: our on-device AI work in a shipping product.
View the case studyNutrify.AI: The Health App Built on Real Science
Nutrify.AI reads blood work and food logs into plain-language guidance, built by DevX and live on iOS and Android.
View the case studyChatFly - AI Communication Platform
ChatFly, an AI communication platform delivered with our partner studio.
View the case studyJoyJoy - AI-Powered Wellness App
JoyJoy, an AI wellness app for iOS delivered with our partner studio.
View the case study
Questions we get asked
How much does AI development cost in San Diego?
- DevX Group LLC bills $95 to $175 per hour depending on the engagement tier, published on our pricing page. A focused first release, one agent or one RAG assistant on your own data, usually takes 4 to 12 weeks. We quote a fixed scope after the week-one discovery so the total is known before build starts.
How long does it take to build an AI agent?
- A prototype scored against real examples takes two to four weeks. A production agent with guardrails, cost controls and a human hand-off path takes 8 to 12 weeks in total. The difference between the two is where most AI projects stall, so we plan for it from day one.
Which AI model will you use?
- The one that scores best on your evaluation set at a cost you can live with. We build on Claude and GPT models through the Vercel AI SDK, so switching models later is a configuration change rather than a rewrite. For privacy-sensitive products we run models locally with Ollama or on-device with Core ML.
Can the AI run without sending our data to a third party?
- Yes. Our own product Parlin runs speech and language models entirely on the Mac, and we build the same way for clients when the data cannot leave the device or the building. Where a cloud model is the right choice, we use API terms that do not train on your data and keep retrieval data in your own Supabase project.
Do you work with companies outside San Diego?
- Yes. DevX Group LLC is based in San Diego and works remotely with clients across the United States. Local clients can meet in person; everyone gets the same bi-weekly demos, milestone reviews and direct access to the engineer writing the code.
Talk to the engineer who will build it
A 30-minute call with Max Sheikhizadeh: what you need, what it costs, and whether we are the right fit.