Showing posts with label Generative AI. Show all posts
Showing posts with label Generative AI. Show all posts

Friday, September 11, 2026

Spec-Driven Development: Building Software from a Clear Specification

See All on GenAI    « Previously

Spec-Driven Development: Building Software from a Clear Specification

A practical guide to SDD, using VS Code and GitHub Copilot, and how it differs from vibe coding.

AI coding assistants have changed the way software is built. Instead of writing every line of code manually, developers can describe what they want and let an AI assistant generate much of the implementation. But there is an important question: what exactly should the AI build?

This is where Spec-Driven Development (SDD) comes in. Rather than starting with code or an informal conversation with an AI assistant, SDD starts with a clear description of the desired behavior, requirements, constraints, interfaces, and acceptance criteria. The specification becomes the source of truth that guides both the developer and the AI coding agent.

The central idea In Spec-Driven Development, you don't primarily tell the AI, "Write some code that does this." You first establish what the software must do, then use the specification to drive the design, implementation, testing, and review.

1. What Is Spec-Driven Development?

Spec-Driven Development is a software development approach in which a written specification acts as the primary guide for creating a feature or system.

The specification describes the intended outcome before implementation begins. Depending on the project, it can contain functional requirements, user stories, API contracts, data models, business rules, constraints, edge cases, non-functional requirements, and acceptance criteria.

The important distinction is that the specification is not merely documentation written after the code. It is an input to the development process.

A simple example

Imagine that you want to add a password-reset feature to an application. A vague instruction might be:

Build a password reset feature.

An AI coding assistant can certainly generate code from this instruction. However, many important questions remain unanswered:

  • How long should a reset link remain valid?
  • Can a reset token be used more than once?
  • What happens when the email address does not exist?
  • How should expired tokens behave?
  • What password rules apply?
  • Should the system reveal whether an email address is registered?
  • What tests must pass before the feature is considered complete?

A specification makes these decisions explicit before implementation.

Feature: Password Reset

Requirements:
1. A user can request a password reset using their email address.
2. The reset token expires after 30 minutes.
3. A token can be used only once.
4. Invalid or expired tokens must be rejected.
5. The system must not reveal whether an email address is registered.
6. The new password must satisfy the application's password policy.

Acceptance criteria:
- Valid reset requests produce a reset email.
- Expired tokens are rejected.
- Reused tokens are rejected.
- Invalid tokens are rejected.
- Successful reset invalidates the token.

Now the implementation has something much more useful than a general instruction: it has a contract to satisfy.

2. Why SDD Matters in AI-Assisted Development

Traditional development already benefits from requirements and design documents. AI-assisted development makes this discipline even more important.

An AI coding assistant can generate code extremely quickly. That is both its strength and its weakness. If the instructions are vague, the assistant can produce a large amount of code that is technically plausible but does not actually solve the intended problem.

SDD addresses this by moving more attention toward the problem definition before asking AI to perform the implementation.

1 Specify Define what must be built.
2 Clarify Resolve ambiguity and edge cases.
3 Plan Decide how the specification will be implemented.
4 Implement Use AI to generate and modify code.
5 Verify Test the implementation against the specification.

Notice that code comes after specification. AI is not removed from the process; instead, its role changes from guessing what you want to implementing a clearly defined requirement.

3. How SDD Can Be Done Using VS Code and GitHub Copilot

VS Code and GitHub Copilot can support an SDD workflow by keeping the specification close to the code and using it as context for AI-assisted planning, implementation, and verification.

The exact Copilot capabilities and interface can evolve over time, but the underlying workflow remains straightforward: write the specification, give the specification to the coding agent, review the plan, implement incrementally, and verify against the requirements.

Step 1: Create a specification

Start by creating a specification file in the repository. The format is less important than the clarity of the content. Markdown is often convenient because it is readable by both humans and AI tools.

docs/specs/password-reset.md

A useful specification should answer questions such as:

  • What problem are we solving?
  • Who is the feature for?
  • What behavior is required?
  • What behavior is explicitly not required?
  • What constraints must be respected?
  • What are the important edge cases?
  • How will we know the implementation is correct?

Step 2: Ask Copilot to analyze the specification

Instead of immediately asking Copilot to write code, ask it to understand and analyze the specification first.

Read docs/specs/password-reset.md.

Analyze the requirements and identify:
- ambiguities
- missing edge cases
- affected parts of the existing codebase
- likely implementation constraints
- tests that will be required

Do not modify any files yet.

This is an important SDD habit: separate understanding from implementation.

Step 3: Ask Copilot for an implementation plan

Once the specification is clear, ask Copilot to create a plan.

Using docs/specs/password-reset.md as the source of truth,
create an implementation plan.

For each step:
- identify the files that need to change
- explain what will change
- identify dependencies
- identify tests that should be added

Do not implement the changes yet.

The plan gives the developer an opportunity to catch incorrect assumptions before code is generated.

Step 4: Review the plan

This is where human judgment becomes particularly important.

Ask questions such as:

  • Does the plan actually satisfy every requirement?
  • Did Copilot misunderstand any business rule?
  • Is the proposed design consistent with the existing architecture?
  • Are security and performance constraints being considered?
  • Are the tests sufficient?

If the plan is wrong, fix the specification or clarify it before implementation begins.

Step 5: Implement in small increments

Instead of giving Copilot one enormous request such as "build the entire feature," work through the plan incrementally.

Implement step 1 from the approved plan.

Use the password-reset specification as the source of truth.
Do not modify unrelated functionality.

After making the changes, explain:
- what changed
- which requirements are now satisfied
- what remains to be implemented

Smaller changes make AI-generated code easier to review, test, and revert.

Step 6: Generate and run tests

Tests should not be an afterthought. The acceptance criteria in the specification should translate into concrete verification.

Review the acceptance criteria in
docs/specs/password-reset.md.

Identify any missing automated tests and add them.
Then run the relevant test suite.

Report:
- tests added
- tests executed
- failures
- requirements that remain unverified

This creates a useful feedback loop:

Specification
      ↓
Implementation plan
      ↓
Code
      ↓
Tests
      ↓
Verification against specification
      ↓
Refinement

Step 7: Keep the specification synchronized

Software changes. Requirements can change too. If the implementation evolves but the specification does not, the repository eventually contains two conflicting sources of truth.

Therefore, when a requirement changes, update the specification deliberately and then update the implementation to match it.

A useful rule: If a behavior matters enough to test or review, it probably matters enough to express clearly in the specification.

4. SDD vs. Vibe Coding

Vibe coding generally describes an AI-assisted development style where the developer gives natural-language instructions, accepts generated code, observes the result, and continues prompting based on what happens.

Vibe coding can be surprisingly productive for prototypes, experiments, throwaway projects, and situations where the cost of being wrong is low.

The problem appears when the same approach is used for software that has important requirements, dependencies, security implications, or long-term maintenance needs.

Aspect Spec-Driven Development Vibe Coding
Starting point A defined specification and acceptance criteria. A natural-language idea or desired outcome.
AI's role Implement and verify against explicit requirements. Generate and refine code through conversational prompts.
Planning Planning is explicit and reviewed before implementation. Planning may emerge during the interaction.
Requirements Made explicit and traceable. Often implicit in the conversation.
Verification Measured against defined acceptance criteria. Often based on whether the result appears to work.
Change management Changes can be reflected in the specification first. Changes may happen through successive prompts.
Best suited for Production systems and requirements-heavy work. Prototypes, exploration, learning, and low-risk experiments.

The key difference

The difference is not simply "writing specifications versus talking to AI." The deeper difference is where the source of truth lives.

In a vibe-coding workflow, the developer's evolving conversation with the AI can become the effective source of truth.

In SDD, the specification is deliberately made explicit so that the developer, AI, reviewer, and tests can all refer to the same intended behavior.

SDD does not mean "never vibe." You can use conversational AI during an SDD workflow. The difference is that the conversation operates within clearly defined requirements rather than replacing them.

5. What to Do in Spec-Driven Development

✓ Do

  • Write requirements before implementation.
  • Define acceptance criteria that can actually be verified.
  • Document important business rules and constraints.
  • Identify edge cases explicitly.
  • Ask Copilot to analyze before asking it to implement.
  • Review AI-generated implementation plans.
  • Implement incrementally.
  • Keep tests aligned with the specification.
  • Review generated code instead of blindly accepting it.
  • Update the specification when requirements change.

✗ Don't

  • Assume AI understands unstated business requirements.
  • Write huge specifications full of unnecessary implementation details.
  • Ask AI to change the entire codebase without boundaries.
  • Treat generated code as automatically correct.
  • Skip tests because the AI says the feature is complete.
  • Allow the specification and implementation to drift apart.
  • Ignore security, performance, or compatibility constraints.
  • Use the specification as an excuse to stop thinking critically.
  • Accept a technically elegant solution that violates the requirements.

6. What Makes a Good Specification?

A good specification should be clear enough to remove important ambiguity without becoming an unnecessarily detailed implementation manual.

For example, instead of specifying every class and method that an AI must create, describe the behavior those components must provide.

Less useful:

Create a PasswordResetService class with a ResetPassword()
method that uses a Dictionary to store tokens.

More useful:

Requirements:
- Generate a cryptographically secure reset token.
- Associate the token with the requesting account.
- Expire the token after 30 minutes.
- Permit a token to be consumed only once.
- Reject invalid and expired tokens.

The second version describes what must be true while leaving room for the implementation to choose an appropriate design.

A practical specification structure

# Feature: Password Reset

## Problem
Users need a secure way to regain access to their account.

## Goal
Allow users to reset their password without administrator intervention.

## Functional Requirements
- ...
- ...
- ...

## Constraints
- ...
- ...
- ...

## Edge Cases
- ...
- ...
- ...

## Acceptance Criteria
- ...
- ...
- ...

## Out of Scope
- ...
- ...

## Verification
- Unit tests
- Integration tests
- Security checks

This structure is simple enough for humans to maintain and structured enough to provide useful context to an AI coding agent.

7. Keep the Specification Focused on Outcomes

One of the easiest mistakes in SDD is turning the specification into a detailed description of the code you have already imagined.

Specifications are generally more valuable when they focus on behavior, requirements, constraints, and outcomes.

For example:

Good:
"The API must return HTTP 404 when the requested product does not exist."

Less useful:
"Create ProductController.GetProduct(), then call ProductRepository.Find()
and return NotFound()."

The first statement defines an externally observable requirement. The second prescribes one particular implementation.

This distinction gives both the developer and AI more flexibility while preserving correctness.

8. SDD Is Not About Writing More Documentation

It is tempting to think that SDD simply means producing more documents. That misses the main point.

The purpose of a specification is to create a shared, explicit contract between the problem, the developer, the AI agent, and the verification process.

A short specification containing ten precise requirements can be more useful than a fifty-page document containing vague prose.

The goal is not documentation for documentation's sake. The goal is to reduce ambiguity.

9. A Practical SDD Workflow for Everyday Development

For a typical feature, you can use the following lightweight workflow in VS Code:

  1. Create a feature specification. Describe the problem, requirements, constraints, edge cases, and acceptance criteria.
  2. Ask Copilot to review it. Have it identify ambiguity, missing cases, and affected parts of the codebase.
  3. Resolve questions. Update the specification until the intended behavior is clear.
  4. Generate an implementation plan. Ask Copilot to propose changes without modifying files.
  5. Review the plan. Confirm that it addresses the specification and fits the existing architecture.
  6. Implement incrementally. Give Copilot one logical part of the approved plan at a time.
  7. Test continuously. Add and run tests derived from the acceptance criteria.
  8. Perform a final specification review. Check that every requirement has been implemented and verified.

This approach preserves one of the biggest advantages of AI coding assistants—speed—while reducing the risk that speed turns into uncontrolled complexity.

10. The Human Developer Still Matters

SDD should not be interpreted as handing the specification to an AI and stepping away.

The developer remains responsible for deciding what the software should do, evaluating trade-offs, understanding the architecture, reviewing the generated implementation, and determining whether the result is safe and correct.

In fact, AI-assisted SDD can make the developer's role more strategic. Instead of spending all of their time typing implementation details, developers can spend more time defining problems, making design decisions, reviewing results, and validating behavior.

Think of the division of labor this way: The human defines the desired outcome and constraints. The AI helps explore and implement solutions. Tests and reviews provide evidence that the implementation actually satisfies the specification.

11. Final Takeaway

Spec-Driven Development is a natural evolution of software development in an age where AI can generate code much faster than humans can manually write it.

The bottleneck increasingly shifts from "How quickly can we write the code?" to "How clearly have we defined what the code should accomplish?"

SDD answers that question by putting the specification at the center of the workflow:

Define the problem
       ↓
Write the specification
       ↓
Clarify requirements
       ↓
Create an implementation plan
       ↓
Review the plan
       ↓
Implement with AI assistance
       ↓
Test
       ↓
Verify against the specification
       ↓
Iterate

Vibe coding can be excellent for discovering possibilities and quickly creating prototypes. But when correctness, maintainability, security, and predictable behavior matter, relying on an evolving conversation alone can become risky.

Spec-Driven Development does not eliminate the speed of AI-assisted coding; it gives that speed direction.

The One-Sentence Definition

Spec-Driven Development is an AI-assisted development approach in which a clear, explicit specification defines the intended behavior first, and the implementation, tests, and review are driven by that specification.


See All on GenAI    « Previously

Sunday, July 5, 2026

Quiz - Fast LLM Inference with Cerebras (Short Course @ DeepLearning.ai)

View Course on DeepLearning.AI    View Other Courses Audited By Us    « Previously





View Course on DeepLearning.AI    View Other Courses Audited By Us    « Previously Tags: Agentic AI,Large Language Models,Generative AI,

Saturday, July 4, 2026

SQLite-Vector: Vector Search in Your Pocket

See All on GenAI    « Previously    Next »

Vector Search in Your Pocket

How sqlite-vector brings AI-powered similarity search to any device — no cloud required

Imagine you have a mobile app that needs to find the most similar image, the best product recommendation, or the right document from a pile of data — all while the user is offline. Traditionally, that would mean sending data to the cloud, running a heavy vector database, and waiting for results. But what if your SQLite database could do all of that, right on the device, with just 30 MB of memory and no indexing wait time? That's exactly what sqlite-vector delivers.

SQLite is already the world's most used database — it's in your phone, your browser, your car, and probably your smart fridge. Sqlite-vector is an extension that adds vector search to SQLite. In plain terms, it lets you store "embeddings" (think of them as mathematical fingerprints of images, text, or audio) and then find the closest matches at lightning speed — all using standard SQL.

Why vector search matters (and why you want it offline)

Modern AI models — from ChatGPT to image recognizers — turn everything into vectors: long lists of numbers that represent the "meaning" of a piece of data. When you want to find something similar, you don't search for exact matches; you search for the nearest neighbors in this high-dimensional space.

Think of it like finding the closest cities on a map — except the map has hundreds of dimensions. That's what vector search does, and it powers:

  • Semantic search — finding documents that are conceptually similar to your query
  • Image retrieval — showing visually similar photos
  • Recommendation systems — matching users with products, videos, or music
  • Voice and audio search — identifying sounds or voice queries
  • Anomaly detection — spotting outliers in sensor data

Until now, doing this on a phone or a low-power device was tricky. You'd need a separate vector database like FAISS or Weaviate, which often means running a server, setting up complex indexes, and waiting hours for preprocessing. Sqlite-vector flips that script.

What makes sqlite-vector different?

Most vector search tools are heavyweight. They require special virtual tables, pre‑indexing phases that can take hours, and external servers. Sqlite-vector takes a radically simpler approach:

✓ Works with ordinary SQLite tables — no special schemas
✓ No preindexing — start searching immediately
✓ Zero‑cost updates — add or change vectors on the fly
✓ Offline first — works without internet
✓ Cross‑platform — iOS, Android, Windows, Linux, macOS
✓ Memory‑efficient — just 30 MB RAM by default

It's built in pure C with SIMD acceleration, which means it runs blazingly fast even on mobile CPUs. And because it's just a SQLite extension, you can drop it into any existing project with minimal effort.

The secret sauce: TurboQuant

One of the coolest features is TurboQuant — a clever quantization technique inspired by a Google Research paper. Instead of storing full-precision vectors (which take up a lot of space), TurboQuant compresses them into 2‑bit, 3‑bit, or 4‑bit representations.

This dramatically reduces memory and storage while still keeping search results accurate. For example, on a dataset of 1 million vectors with 768 dimensions each, raw 32‑bit floats would take about 3 GB. TurboQuant 4‑bit shrinks that to just 396 MB — about 13% of the original size. And the search is still 15 times faster than brute force.

Here's a quick look at the performance on a Mac with ARM64 (NEON):

Mode Quantized storage Full scan / query TurboQuant / query Speedup Recall@10
TurboQuant 4‑bit 396 MB 3248 ms 218 ms 14.9× 0.84
TurboQuant 3‑bit 300 MB 1727 ms 188 ms 9.2× 0.74
TurboQuant 2‑bit 204 MB 3265 ms 85 ms 38.3× 0.48

The 4‑bit mode is a great starting point — it gives a solid balance of speed, memory, and accuracy. For really tight edge budgets, 2‑bit can be a lifesaver, though you'll want to test it with your own data.

Getting started (it's really this simple)

Sqlite-vector is available as a pre‑built binary for all major platforms — Linux, macOS, Windows, Android, and iOS. You can also load it as a WASM module for browsers.

Here's the basic flow in SQL:

-- 1. Load the extension
.load ./vector

-- 2. Create a regular table (no virtual tables needed!)
CREATE TABLE images (
    id INTEGER PRIMARY KEY,
    embedding BLOB,   -- store vectors as binary blobs
    label TEXT
);

-- 3. Insert a vector (as a blob or JSON array)
INSERT INTO images (embedding, label)
VALUES (vector_as_f32('[0.3, 1.0, 0.9, 3.2, ...]'), 'cat');

-- 4. Initialize the vector column
SELECT vector_init('images', 'embedding', 'type=FLOAT32,dimension=384');

-- 5. Quantize for blazing-fast search (TurboQuant 4‑bit)
SELECT vector_quantize('images', 'embedding', 'qtype=TURBO,qbits=4');

-- 6. Search for the top 20 nearest neighbors
SELECT e.id, v.distance
FROM images AS e
JOIN vector_quantize_scan('images', 'embedding', ?, 20) AS v
ON e.id = v.rowid;

That's it. No external servers, no complex indexing, no waiting. Your vector search is ready to go.

💡 Pro tip: You can also use vector_quantize_preload() to load the quantized data into memory for a 4‑5× speedup — perfect for interactive apps.

Where does it shine?

Sqlite-vector is built for Edge AI — scenarios where you need intelligence on the device, not in the cloud.

  • Mobile apps that do on‑device image search, face recognition, or voice commands
  • Privacy‑first applications where data never leaves the user's device
  • Offline‑first tools like note‑taking apps with semantic search
  • Embedded systems in robots, drones, or IoT devices

Because it's a SQLite extension, you also get all the benefits of a full relational database — transactions, joins, filters, and ACID guarantees — combined with vector search.

The bigger picture

Sqlite-vector is part of a larger ecosystem from SQLite AI that's turning SQLite into a complete runtime for intelligent, distributed data. There's also sqlite‑sync for offline‑first sync, sqlite‑ai for on‑device LLM inference, and sqlite‑agent for autonomous AI agents — all living inside your SQLite database.

If you don't want to manage it yourself, SQLite Cloud offers a hosted version with sync, auth, edge functions, and a free tier that gives you 512 MB and 20 connections — no credit card required.

Wrapping up

Sqlite-vector is a game‑changer for anyone building AI‑powered applications that need to work offline, on mobile, or at the edge. It's fast, tiny, and dead simple to use. You don't need to learn a new database or wrestle with complex indexing — just SELECT your way to similar items.

Whether you're building a photo app, a recommendation engine, or a privacy‑first search tool, sqlite‑vector gives you superpowers right inside your SQLite database. And with TurboQuant, you get enterprise‑grade performance on devices that fit in your pocket.

Ready to try it? Head over to the GitHub repository, grab the binary for your platform, and start searching in minutes. The era of on‑device AI is here — and it speaks SQL.

Resources: GitHub · Docs · SQLite AI · Releases

All performance numbers and benchmarks are from the project's official documentation and were measured on macOS ARM64 with the NEON backend. Your mileage may vary depending on hardware and data.

See All on GenAI    « Previously    Next »
Tags: Generative AI,Database

Laptop For Running an LLM Locally (Comparison Specific to India)

See All on GenAI    « Previously    Next »

Laptop For Running an LLM Locally (Comparison Specific to India)

Generated: 2026-07-04
Context: User asked about running LLMs on a Dell Precision with NVIDIA GB10, then explored pricing, shifted to laptops in ₹1‑2 lakh range, upgraded to 12–16 GB VRAM, asked for refurbished MSI Stealth 18, and finally compared a large list of laptops.


1. Initial Query: Dell Precision with NVIDIA GB10

User: what kind of llm can run on this laptop: Dell Pro Precision portfolio with the NVIDIA GB10 Grace Blackwell Superchip

Assistant Clarification:

  • The GB10 is an ARM‑based SoC (NVIDIA Grace CPU + Blackwell GPU)
  • If your Dell Precision has the GB10 chip, it features 128 GB unified memory and 1 PetaFLOP of FP4 AI performance.
  • It can run:
    • Up to ~200B parameter models on a single node (e.g., Llama 3.1 70B, Qwen 2.5 72B, up to 200B with quantization).
    • Two units linked can run 405B models (e.g., Llama 3.1 405B) using 4‑bit quantization.
  • Real‑world speed: ~47 tokens/sec on a 70B coding model.
  • Software: runs NVIDIA DGX OS (Ubuntu for ARM) with PyTorch, Python, Jupyter, and NVIDIA AI Enterprise.

Important: The Intel‑based Dell Precisions with discrete GPUs have far lower VRAM (typically 16–48 GB) and can only run up to ~70B models at best.


2. Pricing of the GB10 System

User: what's it's cost

Starting prices (base model) across regions:

Region Starting Price Currency
USA ~$3,000 – $4,000+ USD
Australia ~10,207.49 AUD
Austria ~6,443.62 EUR
Canada ~8,703.69 CAD
China ~67,270.03 CNY
Hong Kong ~64,365.00 HKD
India ~3,99,000 INR
Ireland ~6,889.75 EUR
New Zealand ~15,328.59 NZD

Note: These are base models. Customisation (storage, memory) increases the price. The GB10 system is specifically the Dell Pro Max line, not all Dell Precisions.


3. More Affordable and Higher‑End Dell Pro Max Options

User asked for cheaper and higher‑end options from the Dell Pro Max line.

Affordable (no GB10, Intel/AMD based)

Model Approx. Price (CNY) Core Specs
Dell Pro Max 14 (AMD) ¥8,999 – ¥13,633 AMD Ryzen AI 5 PRO 340, 32GB
Dell Pro Max Micro (desktop) ¥8,599 – ¥9,252 Intel Core Ultra
Dell Pro Max 16 (laptop) ¥11,999 – ¥24,999 Core Ultra 7/9, up to 64GB, RTX PRO 2000
Dell Pro Max Tower T2 (mid) ¥26,899 – ¥27,899 Core Ultra 9, 64GB, RTX 4080S
Dell Pro Max 14 Premium ~¥21,000+ Core Ultra

High‑End (no GB10, but powerful discrete GPUs)

Model Approx. Price Specs
Dell Pro Max 16 Premium ~$4,583 USD / ¥29,889 Core Ultra 9, RTX PRO 2000
Dell Pro Max 16 Plus ~$6,611 AUD Intel/AMD + RTX PRO
Dell Pro Max Tower T2 (fully loaded) ~$14,349 USD Core Ultra 9 285K, 128GB ECC, RTX PRO 6000 (96GB VRAM)
Dell Precision 7865 Tower ~€3,004 (Ireland) AMD Threadripper PRO, up to 64 cores

Summary: Only the GB10‑based Dell Pro Max (starting ~₹3.99 lakh in India) can run 200B models locally. The tower with RTX PRO 6000 can also handle large models thanks to 96GB VRAM.


4. Laptops in ₹1–2 Lakh for AI Workloads

User: give me some recommendations in 1lakh to 2lakh INR in India for running local AI workloads

Recommended models (all with RTX 4060, 8GB VRAM, ~13B max):

Model Key Specs Approx. Price (INR) Best For
HP Omen 17‑cm2002TX i7‑13700HX / RTX 4060 / 8GB / 16GB ~1,21,999 Best value, large screen
Dell G15 5530 i7‑13650HX / RTX 4060 / 8GB / 16GB ~1,21,990 Reliable Dell build
Acer Nitro V16 Ryzen 7 8845HS / RTX 4060 / 8GB / 16GB ~98,790 Most affordable entry
ASUS ROG Zephyrus G16 Ryzen AI 9 HX 370 / RTX 4060 / 8GB / 32GB ~1,84,990 Premium, portable, OLED
MSI Pulse 16 AI Ultra 7 155H / RTX 4060 / 8GB / 32GB ~1,84,990 High RAM, 140W GPU
Dell Pro 16 Plus (AMD) Ryzen PRO / Integrated NPU (no dGPU) ~96,562 Enterprise, not for LLMs

Key advice: 8GB VRAM is minimum; NPU cannot replace a dGPU. Avoid models without discrete NVIDIA GPUs.


5. Upgrading to 12–16 GB VRAM Laptops

User asked for options with 12‑16 GB VRAM.

Model Specs (CPU/GPU/VRAM/RAM) Approx. Price (INR) Max LLM Size Best For
Dell Alienware m16 Ryzen 9 7845HX / RTX 4080 / 12GB / 32GB 1,49,990 ~30‑40B Best value 12GB
HP Omen 17 i7‑13700HX / RTX 4080 / 12GB / 16GB ~2,39,990 – 2,69,990 ~30‑40B Large screen, powerful
ASUS ROG Zephyrus G16 (2024) Ultra 9 185H / RTX 4080 / 12GB / 32GB 2,76,411 ~30‑40B Premium & portable
ASUS ROG Strix SCAR 16 (2024) i9‑14900HX / RTX 4080 / 12GB / 32GB ~2,77,994 ~30‑40B High‑refresh MiniLED
MSI Raider GE78 HX (14th) i9‑14900HX / RTX 4080 / 12GB / 32GB 2,89,990 ~30‑40B 17‑inch, excellent cooling
MSI Raider A18 HX Ryzen 9 7945HX / RTX 4080 / 12GB / 32GB 2,89,990 ~30‑40B 18‑inch, Wi‑Fi 7
MSI Stealth 18 AI Studio Ultra 9 185H / RTX 4090 / 16GB / 32GB 2,59,990 ~70B Best AI value (16GB)
Razer Blade 18 i9‑14900HX / RTX 4090 / 16GB / 32GB 3,49,990 ~70B Premium build
MSI Raider GE78 HX (4090) i9‑14900HX / RTX 4090 / 16GB / 32GB 4,19,990 ~70B Max 16GB performance

Key: The MSI Stealth 18 AI Studio with RTX 4090 (16GB) is the most affordable way to run 70B‑class models (~₹2.6 lakh). The Alienware m16 (12GB) is the cheapest 12GB option at ₹1.5 lakh.


6. Refurbished MSI Stealth 18 AI Studio in India

User asked: can I get a refurbished model of this at a lower price in India: MSI Stealth 18 AI Studio. Give me some references from Amazon (maybe)

  • New price: RTX 4090 variant: ~₹2,94,990 to ₹3,99,990 depending on retailer. RTX 4080 variant: ~₹3,49,990.
  • Refurbished availability: Not found on Amazon India. International refurbishers (e.g., Reebelo US) list it around $1,658 (~₹1.38 lakh), but shipping + customs + no warranty make it risky.
  • Alternative channels: OLX, Quikr, local computer markets (Nehru Place, Ameerpet).
  • Precautions: Verify in person, check specs, battery health, warranty, and accessories.

Conclusion: Refurbished units are not yet common in India. You may have to wait or buy new.


7. Amazon Options for 12–16 GB VRAM Laptops

User asked to research Amazon and share options.

Model Specs Approx. Price (INR) Max LLM
MSI Stealth 18 AI Studio (RTX 4090) Ultra 9 185H / 16GB / 32GB ~3,99,990 70B
Dell Alienware x16 (RTX 4080) Ultra 9 185H / 12GB / 32GB Price not listed 40B
Dell Alienware m16 R1 (RTX 4080) i9‑13900HX / 12GB / 32GB ~3,94,760 40B
Dell Alienware M18 R2 (RTX 4080) i9‑14900HX / 12GB / 32GB Price not listed 40B
HP OMEN (ck2004TX) RTX 4080 i9‑13900HX / 12GB / 32GB Price not listed 40B
HP Omen 17 (17‑ck2008AX) i7‑13700HX / 12GB / 16GB Price not listed 40B

Observation: Most 4080/4090 laptops are above ₹3 lakh. For a ₹2‑lakh budget, you are limited to 8GB VRAM options.


8. Final Comparison of All Laptops Provided

User dumped a long list of laptops and asked to compare with specific columns.

Here is the comprehensive comparison table with all models ranked by VRAM and capability:

Laptop Model Key Specs (CPU / GPU / VRAM / RAM) Approx. Price (INR) MAX LLMs Size Supported OS Support Best For
Alienware 16 Area‑51 Ultra 9‑275HX / RTX 5090 / 24GB / 64GB ₹4,84,990 ~70B+ Win 11 + MSO Maximum AI performance
Lenovo Legion Pro 7 2025 Ultra 9‑275HX / RTX 5090 / 24GB / 64GB ₹4,72,490 ~70B+ Win 11 + Office Best value for 24GB VRAM
ASUS ProArt P16 OLED (2025) Ryzen AI 9 HX 370 / RTX 5090 / 24GB / 64GB ₹4,19,990 ~70B+ Win 11 + M365 Creator-focused, 4K OLED
NXTGN XP4 (Desktop) i9‑14900K / RTX 5060 Ti / 16GB / 64GB Not listed ~70B Win 11 Pro Desktop upgradeability
Lenovo Legion 9 i9‑13980HX / RTX 4090 / 16GB / 32GB ₹4,49,510 ~70B Win 11 + Office Previous‑gen flagship
ASUS ROG Strix SCAR 16 Ultra 9‑275HX / RTX 5080 / 16GB / 32GB ₹3,79,990 ~70B Win 11 + M365 High‑end gaming & AI
ASUS Zenbook 14 (2026) Ultra 9‑285H / Integrated iGPU / shared / 32GB ₹1,19,990 Light AI only Win 11 + M365 Ultra‑portable productivity
ASUS Zenbook S16 (2026) Ryzen AI 9‑465 / Integrated iGPU / shared / 32GB ₹1,69,990 Light AI only Win 11 + M365 Premium ultraportable
Lenovo Yoga Slim 7 Ultra 7‑155H / Integrated iGPU / shared / 16GB ~₹90,990 Light AI only Win 11 + Office Affordable ultraportable
Lenovo IdeaPad Slim 5 Ultra 7‑355 / Integrated iGPU / shared / 16GB ₹1,36,990 Light AI only Win 11 + MSO Next‑gen AI PC (Copilot+)
MSI Stealth 16 AI Studio Ultra 9‑185H / RTX 4070 / 8GB / 32GB ₹2,43,990 ~13B Win 11 Pro Slim & portable AI
ASUS ROG Strix G16 Ultra 9‑275HX / RTX 5070 / 8GB / 32GB ₹2,59,990 ~13B Win 11 + M365 Gaming & AI entry
HP Omen 16 Max Ryzen AI 9 HX 375 / RTX 5070 Ti / 12GB / 32GB ₹2,67,990 ~40B Win 11 + Office Best mid‑range AI value
HP Omen (an0015TX) Ultra 7‑255H / RTX 5060 / 8GB / 24GB ₹1,51,490 ~13B Win 11 + M365 Budget gaming & AI
HP Victus (fa2382tx) i5‑14450HX / RTX 4050 / 6GB / 24GB ₹1,02,990 ~7B Win 11 + M365 Most affordable option
HP Victus (fa2531TX) i7‑13650HX / RTX 4050 / 6GB / 24GB ₹1,19,990 ~7B Win 11 + M365 Budget AI with better CPU

9. Final Recommendation (Based on Budget)

If Budget ~₹4‑5 Lakh (for 70B models)

  • Best Value: Lenovo Legion Pro 7 2025 (₹4.72L) – similar to Alienware but cheaper.
  • Best for Creators: ASUS ProArt P16 OLED (₹4.19L) – lighter, superior display.
  • Best Performance: Alienware 16 Area‑51 (₹4.84L) – top‑tier cooling and build.

If Budget ~₹2.5‑3 Lakh (best balance)

  • HP Omen 16 Max (₹2.67L) – 12GB RTX 5070 Ti allows up to 40B models – the sweet spot.

If Budget ₹1‑2 Lakh (entry‑level)

  • HP Omen (an0015TX) (₹1.51L) – 8GB RTX 5060, runs up to 13B.
  • HP Victus (fa2382tx) (₹1.02L) – 6GB RTX 4050, runs up to 7B.

Avoid these for LLMs

  • All laptops with only integrated graphics (Intel Arc, AMD Radeon, or NPU) – they cannot run models larger than a few billion parameters efficiently.

10. Critical Takeaways to Avoid Wrong Decisions

  1. VRAM is the absolute king – more VRAM = larger models you can run locally.
  2. System RAM does NOT substitute VRAM – the model must fit into GPU memory.
  3. NPUs (AI accelerators) are for light, energy‑efficient tasks – they cannot run large LLMs (7B+).
  4. Price vs. capability:
    • Up to 13B → 8GB VRAM (₹1‑1.8L)
    • Up to 40B → 12GB VRAM (₹2.6‑3L)
    • Up to 70B → 16‑24GB VRAM (₹3.8‑4.8L)
  5. Refurbished high‑end laptops are rare in India – be prepared to buy new or wait.
  6. Always verify GPU model and VRAM before purchase – don't rely on “AI PC” marketing.
  7. Linux compatibility is generally good for NVIDIA GPUs – but integrated NPUs may have limited driver support.

See All on GenAI    « Previously    Next »

Thursday, June 18, 2026

ROUGE Score (Evaluating Machine Translation and Other GenAI Use Cases)

See All on GenAI    « Previously    Next »
You may recall that BLEU score was precision oriented. Next, we will see a Recall oriented metric to evaluate Machine Translation and other GenAI use cases.

See All on GenAI    « Previously    Next »