Back to Home

Under the Hood

A technical deep dive into GeoViT: the dataset, the Trinity architecture, ensemble learning, IMINT micro-clues, experiments and optimization.

Elber Dalfidan, Senior Software Developer

Introduction

A comprehensive technical journey through GeoViT’s architecture, from project inception to cutting-edge ensemble learning and forensic image intelligence. GeoViT is an R&D project at nadcode.

Training Images Istanbul Districts Peak Accuracy Inference Time
66,340 39 81-83% <1s

What You’ll Learn

  • How we built a 4-layer hierarchical architecture that combines zero-shot learning, OCR, and deep learning
  • The evolution from single-model (75% accuracy) to 4-model ensemble (81-83% accuracy)
  • Memory-efficient sequential loading strategy for M3 MacBook Air (16GB)
  • Sherlock IMINT: Forensic micro-clue analysis (vehicles, flora, infrastructure, signage)
  • Research experiments that worked (and those that didn’t) with honest lessons learned
  • Complete training pipeline: data collection, geographic grid system, and geo-loss function

Project Timeline

The evolution of GeoViT through three major development phases.

  1. Phase 1: Foundation

    January 22, 2026 · 831a6aa

    Initial prototype with Trinity Architecture and geographic grid system

    Key Achievements

    • Built 4-layer Trinity cascade (CLIP → OCR → Visual → Oracle)
    • Implemented 2000-cell geographic grid across Istanbul
    • Achieved baseline 75% accuracy with single ViT model
    • Integrated CLIP for landmark detection, EasyOCR for Turkish text
    Accuracy
    75%
    Grid Cells
    2000
    Inference
    200ms
  2. Phase 2: Ensemble Enhancement

    January 31, 2026 · b1eec3a

    4-model ensemble with sequential loading, TTA, and post-processing

    Key Achievements

    • Sequential ensemble: GeoViT (40%), DeiT (25%), ConvNeXt (20%), EfficientNet (15%)
    • Memory-efficient loading for M3 16GB (peak 1.5GB)
    • Test-Time Augmentation with 8 street-view optimized transforms
    • Spatial smoothing + coordinate refinement post-processing
    • Accuracy boost: 75% → 81-83% (+6-8 points)
    Accuracy
    81-83%
    Models
    4
    Inference
    1-3s
  3. Phase 3: Sherlock IMINT

    Current · In Progress

    Forensic micro-clue analysis for low-confidence predictions

    Key Achievements

    • IMINT (Imagery Intelligence) extracts 4 micro-clue categories
    • Vehicles: Turkish plates, yellow taxis, İETT buses, Metrobüs
    • Flora: Mediterranean vegetation, coastal vs inland patterns
    • Infrastructure: European vs Asian architecture, streetlight styles
    • Signage: District logos, Metro markers, directional signs
    • Ensemble voting (3 passes) for robust clue detection
    Accuracy Boost
    +2%
    Clue Types
    4
    Latency
    +3.5s

The Challenge

Visual geolocation in Istanbul presented four critical challenges that shaped our architectural decisions.

Challenge 1: Ambiguous Visual Features

Generic street scenes lack distinctiveness for accurate geolocation.

  • Similar apartment buildings across multiple districts
  • Generic street furniture (benches, lampposts) provides weak signals
  • Visual-only models plateau at ~75% accuracy on ambiguous scenes
  • Need for ensemble diversity to capture different feature types

Example: A residential street with standard apartment buildings could exist in Kadıköy, Şişli, or Üsküdar. Without distinct landmarks, visual features alone cannot differentiate.

Challenge 2: Ignoring Text Signals

Vision models treat text as visual patterns, missing semantic ground truth.

  • Street signs contain explicit district names (e.g., “Kadıköy Belediyesi”)
  • ViT models see text as pixel patterns, not semantic information
  • Turkish characters (ı, ğ, ş) require specialized OCR
  • Text signals provide 100% ground truth when readable

Solution: OCR Notary layer extracts Turkish text with EasyOCR and overrides visual predictions when district keywords are found with >70% confidence.

Challenge 3: Memory Constraints

Target hardware (M3 MacBook Air 16GB) cannot hold 4 models simultaneously.

  • 4 models × ~330MB each = 1.3GB base memory
  • Activation tensors during inference add ~500MB per model
  • Parallel loading causes OOM errors on 16GB systems
  • Need sequential loading: load → predict → unload → next model

Solution: Sequential loading with MPS cache clearing keeps peak memory at ~1.5GB.

for model in [GeoViT, DeiT, ConvNeXt, EfficientNet]:
    load_to_mps() → predict() → del model → torch.mps.empty_cache()

Challenge 4: Confidence Calibration

Low-confidence predictions still make guesses instead of asking for help.

  • Softmax probabilities are often overconfident (calibration problem)
  • 60% confidence on ambiguous scene is unreliable
  • Need intelligent fallback when visual models are uncertain
  • When to trust vision vs invoke LLM reasoning?

Solution: Oracle layer (Claude/Gemini) activates when ensemble confidence <60%, enriched with IMINT micro-clue context for better reasoning.

Why Istanbul?

Geographic Diversity

  • 39 distinct administrative districts
  • European vs Asian sides (Bosphorus divide)
  • Coastal, urban, suburban, and industrial zones
  • Elevation ranges from sea level to 500m hills

Architectural Diversity

  • Historic districts (Fatih, Beyoğlu) with Ottoman architecture
  • Modern business districts (Maslak, Levent) with skyscrapers
  • Residential areas with mixed building styles
  • Industrial zones (Tuzla, Pendik) with distinct features

Data Availability

  • Rich Mapillary coverage (66,340 street-view images)
  • High-quality geotagged data
  • Comprehensive district coverage
  • Temporal diversity (seasons, time of day)

Perfect Test Bed

  • Complex enough to be challenging
  • Manageable scope (single city)
  • Turkish language OCR requirements
  • Real-world application potential

Dataset & Training

How we built and processed 66,340 street-view images into a production-ready geolocation dataset.

Data Sources

Mapillary API

Primary data source for geotagged street-view imagery across Istanbul.

  • 66,340 street-view images downloaded via API
  • Each image tagged with precise latitude/longitude coordinates
  • Coverage across all 39 Istanbul districts
  • Temporal diversity: different seasons, times of day, weather conditions

Geographic Grid System

Instead of regression (predicting continuous lat/lon), we use classification with a 2000-cell grid. This approach provides better generalization and more stable training.

Grid Specifications

  • 2000 cells across Istanbul
  • Cell size: ~150m × 150m
  • Bounding box: (28.4°W, 40.75°S, 29.70°E, 41.30°N)
  • Covers all 39 districts comprehensively

Why Classification?

  • More stable training than regression
  • Better generalization on test set
  • Easier to interpret model confidence
  • Allows spatial smoothing post-processing

Conversion Process

Raw Image (lat=41.0082, lon=28.9784)
  ↓ Geocoding
Grid Cell ID: 1247 (Beşiktaş district)
  ↓ Training Target
One-hot vector [0, 0, ..., 1, ..., 0] (2000 dims)
  ↓ Inference
Predicted Cell: 1247 → Convert to (41.008, 28.978)

Data Processing Pipeline

  1. Raw Images (Mapillary): Download via API with geotags
  2. Geocoding: Convert lat/lon → grid cell ID using bounding box math
  3. Train/Val/Test Split: 80/10/10 stratified by district for balanced representation
  4. Augmentation: Flip, rotate, color jitter, random crop during training
  5. ViT Preprocessing: Resize to 224×224, normalize with ImageNet stats

Training Process

Training Configuration

  • Loss function: Geo-loss (classification + distance penalty)
  • Optimizer: AdamW with cosine schedule
  • Batch size: 32 (M3 memory constraint)
  • Epochs: 15-20 until convergence
  • Hardware: M3 MacBook Air (MPS backend)

Geo-Loss Function

Combines classification loss with geographic distance penalty to prioritize nearby cells over distant ones.

loss = cross_entropy(pred, true_cell)
  + α * distance_penalty(pred, true_coords)

# Prefer wrong nearby cells
# over wrong distant cells

Training Command

python3 -m src.training.train_v4_geoloss \
  --data_dir data/processed/istanbul \
  --epochs 15 \
  --batch_size 32 \
  --lr 1e-4 \
  --model_name geovit_v4

Dataset Statistics

Metric Value Notes
Total Images 66,340 80/10/10 train/val/test split, stratified by district
Districts Covered 39 Balanced representation across all districts
Grid Cells 2000 ~150m × 150m cell resolution

Trinity Architecture

A 4-layer hierarchical cascade where each layer can override the next, combining zero-shot learning, OCR, deep learning, and LLM reasoning.

Core Philosophy

The Trinity Architecture uses a hierarchical override cascade where higher layers can short-circuit lower layers based on confidence thresholds. This ensures fast inference on easy cases while falling back to more expensive reasoning when needed.

Why “Trinity”? Originally referred to the core 3 layers (CLIP, OCR, Visual). The Oracle was added later as a fourth fallback layer for low-confidence predictions.

Decision Flow

  1. CLIP Guard

    Landmark detected? Threshold: 90%

    Return landmark location

    Pass to OCR

  2. OCR Notary

    District text found? Threshold: 70%

    Return text-based location

    Pass to Ensemble Expert

  3. Ensemble Expert

    Visual confidence high? Threshold: 60%

    Return visual prediction

    Pass to Oracle

  4. Oracle

    LLM reasoning Threshold: Fallback

    Return LLM prediction

Cascade Logic

Each layer processes the image in order. If a layer meets its confidence threshold, it returns immediately. Otherwise, the image cascades to the next layer. The Oracle is the final fallback for low-confidence predictions.

Layer 1: CLIP Guard

Zero-shot landmark detection for famous Istanbul locations.

How It Works

  • Model: OpenAI CLIP (ViT-B/32)
  • Method: Text-image similarity scoring
  • Threshold: 90% confidence
  • Inference: ~0.1s
  • Landmarks: 12 famous locations

Examples

  • Galata Tower (Beyoğlu)
  • Hagia Sophia (Fatih)
  • Blue Mosque (Fatih)
  • Maiden’s Tower (Üsküdar)
  • Bosphorus Bridge (Beşiktaş/Üsküdar)
  • Dolmabahçe Palace (Beşiktaş)

Code Snippet

# CLIP Guard inference loop
for landmark_name, config in ISTANBUL_LANDMARKS.items():
    prompts = config["prompts"]  # e.g., ["a photo of Galata Tower"]
    similarity = clip_model(image, prompts)

    if similarity > 0.90:  # High confidence threshold
        return {
            "district": config["district"],
            "coordinates": config["coordinates"],
            "confidence": similarity,
            "source": "CLIP Guard"
        }

Limitation: Only works for famous landmarks with distinctive visual features. Generic street scenes pass through to next layer.

Layer 2: OCR Notary

Extract Turkish text from signs and match district keywords.

How It Works

  • Model: EasyOCR (Turkish + English)
  • Method: Text extraction + keyword matching
  • Threshold: 70% OCR confidence
  • Inference: ~0.3s
  • Keywords: District names, Metro stations

Examples

  • “Kadıköy Belediyesi” → Kadıköy
  • “Beşiktaş Metro” → Beşiktaş
  • “Şişli AVM” → Şişli
  • “Üsküdar İskele” → Üsküdar
  • “Fatih Municipality” → Fatih

Code Snippet

# OCR Notary text extraction
results = ocr_reader.readtext(image_array, paragraph=False)

for (bbox, text, confidence) in results:
    if confidence < 0.70:  # Skip low-confidence detections
        continue

    text_clean = text.lower().strip()

    # Match against district keywords
    for district, keywords in DISTRICT_KEYWORDS.items():
        if any(kw in text_clean for kw in keywords):
            return {
                "district": district,
                "confidence": confidence,
                "matched_text": text,
                "source": "OCR Notary"
            }

Limitation: Requires readable text in the image. Blurry, occluded, or non-text images pass through to next layer.

Layer 3: Ensemble Expert

Visual pattern recognition with ensemble learning (optional).

Single Model Mode

  • Model: GeoViT v4 (ViT-base)
  • Training: Geo-loss on 66,340 images
  • Threshold: 60% confidence
  • Inference: ~0.2s
  • Accuracy: 75%

Ensemble Mode

  • Models: GeoViT, DeiT, ConvNeXt, EfficientNet
  • Strategy: Sequential loading + weighted voting
  • Inference: ~1-3s (with TTA)
  • Accuracy: 81-83%
  • Memory: Peak 1.5GB

Code Snippet (Ensemble)

# Sequential ensemble loading (memory-efficient)
model_configs = [
    {"name": "geovit_v4", "weight": 0.40},
    {"name": "deit_base", "weight": 0.25},
    {"name": "convnext_tiny", "weight": 0.20},
    {"name": "efficientnet_v2_s", "weight": 0.15}
]

accumulated_probs = torch.zeros(2000)  # 2000 grid cells

for config in model_configs:
    model = load_model(config["name"]).to("mps")
    probs = model(image_tensor)
    accumulated_probs += config["weight"] * probs

    del model  # Unload from memory
    torch.mps.empty_cache()  # Clear MPS cache

final_prediction = accumulated_probs.argmax()

Limitation: Ambiguous scenes with low visual distinctiveness (<60% confidence) trigger the Oracle layer.

Layer 4: Oracle

LLM-based reasoning with IMINT context (low-confidence fallback).

How It Works

  • Models: Claude CLI or Gemini 2.0
  • Trigger: Ensemble confidence <60%
  • Context: IMINT micro-clues (see Sherlock IMINT)
  • Inference: ~3.5s
  • Use case: Ambiguous residential streets

IMINT Integration

  • Extracts vehicles (taxis, buses, plates)
  • Analyzes flora (Mediterranean vegetation)
  • Detects infrastructure (architecture style)
  • Reads signage (district logos, Metro)
  • Provides structured context to LLM

Code Snippet

# Oracle invocation with IMINT context
if ensemble_confidence < 0.60:
    # Extract micro-clues
    imint_context = analyze_micro_clues(image_path)

    # Build prompt with context
    prompt = f"""
    Analyze this Istanbul street-view image.

    Ensemble prediction: {ensemble_district} ({ensemble_confidence:.1%})

    IMINT Context:
    - Vehicles: {imint_context['vehicles']}
    - Flora: {imint_context['flora']}
    - Infrastructure: {imint_context['infrastructure']}
    - Signage: {imint_context['signage']}

    Which Istanbul district is this most likely?
    """

    oracle_result = llm_client.invoke(prompt, image=image_path)
    return parse_oracle_response(oracle_result)

Performance: Oracle adds ~3.5s latency but improves accuracy by ~2% on low-confidence predictions. Only activated for ~15% of images.

Ensemble Evolution

From single-model baseline to 4-model ensemble with sequential loading, TTA, and post-processing.

Why Ensemble?

Single models plateau at 75% accuracy because they learn similar features. Ensemble diversity allows each model to capture different aspects of the scene.

The Problem

  • Single GeoViT v4: 75% accuracy ceiling
  • High confidence on wrong predictions
  • Struggles with ambiguous residential streets
  • Similar errors across training runs

The Solution

  • 4 diverse models with different architectures
  • Weighted soft voting combines strengths
  • Ensemble: 81-83% accuracy (+6-8 points)
  • More robust confidence calibration

Model Selection

Model Weight Description Model Size Key Strength
GeoViT v4 40% Anchor model trained with geo-loss. Best single-model performance. ~330 MB Fine-tuned on Istanbul, strong on landmarks
DeiT-Base 25% Distilled Vision Transformer. Efficient and generalizes well. ~330 MB Good on generic scenes, robust to variations
ConvNeXt-Tiny 20% Modernized ConvNet. Strong local pattern recognition. ~110 MB Captures architectural details, street patterns
EfficientNetV2-S 15% Mobile-optimized CNN. Fast and memory-efficient. ~85 MB Good on texture patterns, building materials

Why These 4 Models?

Selected through ablation studies. Each model contributes unique features:

  • Architecture diversity: ViT (GeoViT, DeiT) + CNN (ConvNeXt, EfficientNet)
  • Scale diversity: Different model sizes capture features at different scales
  • Training diversity: Different initialization and augmentation strategies
  • Complementary errors: Models fail on different subsets of images

Sequential Loading Strategy

The Memory Challenge

Target hardware: M3 MacBook Air with 16GB unified memory. Loading all 4 models simultaneously causes OOM (Out of Memory) errors.

# Parallel loading (FAILS on M3 16GB)
models = [
    GeoViT().to("mps"),      # ~330MB + 400MB activations
    DeiT().to("mps"),        # ~330MB + 400MB activations
    ConvNeXt().to("mps"),    # ~110MB + 300MB activations
    EfficientNet().to("mps") # ~85MB + 250MB activations
]
# Total: ~2.2GB - causes OOM with other processes

# Sequential loading (WORKS on M3 16GB)
for model_config in ensemble_models:
    model = load_model(config).to("mps")  # ~800MB peak
    probs = model(image)
    accumulated += config["weight"] * probs
    del model
    torch.mps.empty_cache()  # Release memory
# Peak: ~1.5GB - fits comfortably
Peak Memory Speed Fits M3 16GB
1.5GB 4× slower (vs parallel) Yes

Test-Time Augmentation (TTA)

Apply 8 street-view optimized augmentations at inference time and average predictions for improved robustness.

Augmentation Purpose
Horizontal Flip Mimic reverse driving direction
Rotate ±5° Camera angle variations
Random Crop Different framing
Brightness ±20% Time of day variations
Contrast ±15% Weather conditions
Saturation ±10% Color calibration
Hue ±5° White balance shifts
Gaussian Blur Motion blur simulation

TTA Strategy

# Average predictions across augmentations
tta_probs = []
for augmentation in [flip, rotate, crop, brightness, ...]:
    aug_image = augmentation(original_image)
    probs = ensemble_predict(aug_image)
    tta_probs.append(probs)

final_probs = torch.stack(tta_probs).mean(dim=0)  # Average
prediction = final_probs.argmax()
  • Benefit: +2.5% accuracy improvement on test set
  • Cost: 8× inference time (3s total with ensemble)

Post-Processing Pipeline

Spatial Smoothing

Blend probabilities of neighbor cells (α=0.15) to create smoother predictions.

smoothed[cell] = (1-α) * pred[cell] + α * mean(neighbors)

Impact: +0.5% accuracy

Coordinate Refinement

Weighted centroid of top-10 cells instead of single cell center.

coords = Σ(top_k_cells * their_probs) / Σ(top_k_probs)

Impact: +0.3% accuracy

Temperature Scaling

Calibrate confidence scores with T=1.4 (found via grid search).

calibrated_probs = softmax(logits / T)

Impact: Better uncertainty estimates

Performance Comparison

Configuration Accuracy Inference Time Use Case
Single GeoViT v4 75.0% 200ms Real-time apps
Ensemble (no TTA) 78.5% 1000ms Production (recommended)
Ensemble + TTA 81.0% 3000ms Batch processing
Ensemble + TTA + PostP 81.5% 3200ms Maximum accuracy
+ IMINT (low conf) 83.0% 6700ms Critical accuracy

Sherlock IMINT

Forensic imagery intelligence that extracts micro-clues invisible to ML models.

What is IMINT?

IMINT (Imagery Intelligence) is a military and forensic technique for extracting actionable information from images. We adapt this methodology to visual geolocation by analyzing micro-clues that ML models overlook.

Traditional Approach (Vision Models)

  • Learn pixel patterns end-to-end
  • Miss semantic details (license plate text)
  • No explicit reasoning process
  • Low confidence on ambiguous scenes

IMINT Approach (Sherlock)

  • Extract structured micro-clues
  • Semantic understanding (34 = Istanbul plate)
  • Explicit reasoning chain
  • Provides context to Oracle LLM

Why Sherlock IMINT?

The Problem

When ensemble confidence is low (<60%), we invoke the Oracle LLM. But generic “analyze this image” prompts waste the LLM’s reasoning capacity on basic visual tasks.

The Solution

Sherlock IMINT pre-processes the image to extract structured micro-clues, then provides this context to the Oracle. This focuses the LLM’s attention on high-level reasoning rather than low-level visual parsing.

Impact

Accuracy Boost Latency Added Images Triggered
+2% +3.5s ~15%

The 4 Micro-Clue Categories

Key indicators for geolocation, grouped into four categories.

1. Vehicles

  • Turkish license plates (34 prefix = Istanbul)
  • Yellow taxis (ubiquitous in Istanbul)
  • İETT buses (public transport branding)
  • Metrobüs (BRT system, specific routes)
  • Dolmuş minibuses (district-specific colors)

Analysis prompt:

Look for vehicles, license plates (34 prefix?), taxis (yellow?), buses (İETT branding?), Metrobüs.

2. Flora

  • Mediterranean vegetation patterns
  • Coastal palm trees (Bosphorus areas)
  • Plane trees (common in European side)
  • Pine trees (Asian side hills)
  • Seasonal indicators (spring blossoms)

Analysis prompt:

Analyze vegetation: Mediterranean species? Palm trees (coastal)? Pine trees (Asian side)? Plane trees (European side)?

3. Infrastructure

  • European vs Asian side architecture
  • Modern (Maslak/Levent skyscrapers)
  • Historic (Fatih/Beyoğlu Ottoman buildings)
  • Streetlight styles (vary by district)
  • Pavement patterns and materials
  • Power line configurations

Analysis prompt:

Examine architecture: Modern vs historic? European vs Asian style? Streetlight design? Pavement type?

4. Signage

  • Blue directional signs (Istanbul standard)
  • Municipality logos (district branding)
  • Metro/tramvay station markers
  • AVM (shopping mall) names
  • District boundary signs

Analysis prompt:

Read all visible text: District names? Municipality logos? Metro stations? Street signs? Mall names?

Implementation

Sherlock IMINT uses ensemble voting (3 passes with Claude CLI) to extract robust micro-clues, then provides structured context to the Oracle.

def analyze_micro_clues(image_path: str) -> dict:
    """Extract IMINT micro-clues via ensemble voting."""

    clue_categories = ["vehicles", "flora", "infrastructure", "signage"]
    results = {cat: [] for cat in clue_categories}

    # Run 3 independent passes for robustness
    for pass_idx in range(3):
        for category in clue_categories:
            prompt = IMINT_PROMPTS[category]  # Category-specific prompt
            response = claude_cli.invoke(prompt, image=image_path)
            parsed = parse_imint_response(response, category)
            results[category].append(parsed)

    # Majority voting
    consensus = {}
    for category in clue_categories:
        # Find most common answer across 3 passes
        consensus[category] = majority_vote(results[category])

    return consensus

# Example output:
# {
#   "vehicles": "Yellow taxis, 34 license plates, İETT bus visible",
#   "flora": "Mediterranean plane trees, coastal vegetation",
#   "infrastructure": "Modern streetlights, European-style pavement",
#   "signage": "Beşiktaş Municipality logo, Metro sign"
# }

Ensemble Voting Strategy

  • Run analysis 3 times independently
  • Majority vote on each clue category
  • Filters out hallucinations and noise
  • More robust than single-pass

Oracle Integration

  • IMINT context provided to Oracle LLM
  • Structured JSON format for parsing
  • LLM focuses on high-level reasoning
  • Improves accuracy on ambiguous cases

Performance Impact

Metric Value Notes
Accuracy Boost +2% On low-confidence predictions where Oracle is invoked
Latency Added +3.5s 3 Claude CLI passes + parsing overhead
Images Triggered ~15% Only activated when ensemble confidence <60%

Example Scenario

Image: Generic residential street with no landmarks.

  • Before: Ensemble: Kadıköy (58% confidence) → Oracle invoked → Generic analysis → Wrong: Şişli
  • After: IMINT: “Yellow taxi, 34 plate, İETT bus, European architecture” → Oracle with context → Correct: Kadıköy

What Worked & What Didn’t

Honest lessons from research experiments, including failures and pivots.

Experiments That Worked

Geographic Loss Function

  • Problem: Standard cross-entropy treats all wrong predictions equally
  • Solution: Geo-loss penalizes distant predictions more than nearby ones
  • Result: +3% accuracy improvement
loss = cross_entropy(pred, true_cell) + α * distance_penalty(pred_coords, true_coords)

Sequential Ensemble Loading

  • Problem: Parallel loading causes OOM on M3 16GB
  • Solution: Load → predict → unload → next model
  • Result: Memory: 3GB → 1.5GB (fits M3)
for model in models:
  load() → predict() → del model → torch.mps.empty_cache()

OCR Override Strategy

  • Problem: Vision models ignore text semantics
  • Solution: Trust OCR text when confidence >70%, override visual predictions
  • Result: +5% accuracy on text-heavy images
if ocr_confidence > 0.70:
  return ocr_district  # Override visual prediction

Temperature Scaling

  • Problem: Softmax probabilities are overconfident
  • Solution: Calibrate with T=1.4 (found via grid search)
  • Result: Better uncertainty estimates, more reliable confidence scores
calibrated_probs = softmax(logits / 1.4)

Experiments That Failed

Regression-Based Approach

  • What we tried: Direct lat/lon prediction instead of classification
  • The problem: Poor generalization, high variance in predictions
  • Why it failed: Continuous space too large, model struggles to converge
  • Switched to: Classification with 2000-cell grid → stable training

Parallel Ensemble Inference

  • What we tried: Load all 4 models simultaneously for faster inference
  • The problem: OOM (Out of Memory) errors on M3 16GB
  • Why it failed: 4 models × 330MB + activations > 2GB memory
  • Switched to: Sequential loading → fits memory, 4× slower but works

Random TTA Transforms

  • What we tried: Apply generic augmentations (vertical flip, extreme rotation)
  • The problem: Degraded accuracy by -1%
  • Why it failed: Street-view has specific geometry (no upside-down scenes)
  • Switched to: Curated 8 street-view specific transforms

Attention Pooling

  • What we tried: Attention-based ensemble aggregation instead of voting
  • The problem: Slower inference, no accuracy gain
  • Why it failed: Overcomplicates the problem, simple voting works well
  • Switched to: Weighted soft voting → simpler and just as good

Lessons Learned

  1. Simplicity often wins: Weighted soft voting outperforms complex attention mechanisms
  2. Text signals are gold: OCR override provides 100% ground truth when readable
  3. Memory constraints drive architecture: Target hardware shapes design decisions
  4. Domain-specific augmentations matter: Generic TTA can hurt more than help
  5. Classification > Regression: For geolocation, discrete cells generalize better than continuous coords
  6. Failed experiments are valuable: Each failure taught us what NOT to do

Performance Optimization

Balancing accuracy, speed, and memory constraints for production deployment.

Memory Optimization

Target hardware: M3 MacBook Air with 16GB unified memory. All optimizations focus on fitting within this constraint.

Techniques

  • Sequential loading: One model at a time
  • MPS backend: Metal Performance Shaders on M3
  • Batch size tuning: 32 optimal for training
  • Gradient checkpointing: During training
  • Cache clearing: torch.mps.empty_cache()

Memory Profile

Configuration Peak Memory
Single Model ~800MB
Ensemble (Sequential) ~1.5GB
Ensemble (Parallel, OOM) ~3.2GB

Memory Management Code

# Proper memory management for MPS
import torch

for model_config in ensemble_models:
    # Load model to MPS device
    model = load_model(config["name"]).to("mps")

    # Inference
    with torch.no_grad():
        probs = model(image_tensor)

    # Accumulate weighted probabilities
    accumulated_probs += config["weight"] * probs

    # Critical: Clean up immediately
    del model
    torch.mps.empty_cache()  # Release Metal memory

# Peak memory stays under 1.5GB

Speed Optimization

Early Exit Strategies

Trinity layers short-circuit when high confidence is achieved, avoiding expensive inference.

  • CLIP Guard: ~0.1s (if landmark detected)
  • OCR Notary: ~0.3s (if text found)
  • Ensemble: ~1-3s (most images)
  • Oracle: +3.5s (only ~15% of images)

Caching Strategies

Pre-compute and cache static resources to avoid redundant work.

  • CLIP embeddings: Cache for 12 landmarks
  • OCR keywords: Pre-compiled regex
  • Grid coordinates: Pre-computed lookup table
  • Model weights: Load once, reuse

Latency Breakdown (Average Image)

Task Time
CLIP Guard check 100ms
OCR Notary extraction 300ms
Ensemble inference (4 models) 1000ms
Post-processing 50ms
Coordinate conversion 10ms

Accuracy vs Speed Tradeoffs

Mode Accuracy Latency Use Case Recommended
Single GeoViT 75% 200ms Real-time apps, demos
Ensemble (no TTA) 78.5% 1000ms Production deployment Best
Ensemble + TTA 81.5% 3000ms Batch processing
+ IMINT (low conf) 83% 6500ms Critical accuracy needs

Recommendation: Ensemble (no TTA) provides the best accuracy/latency tradeoff for production. TTA and IMINT can be enabled selectively for high-value predictions.

Device Compatibility

Device Status Description Inference Time
MPS (M1/M2/M3) Optimized Primary target. Metal Performance Shaders on Apple Silicon. 200ms (single) / 1000ms (ensemble)
CUDA (NVIDIA GPU) Supported Supported for users with NVIDIA GPUs. 150ms (single) / 800ms (ensemble)
CPU Slow Fallback mode when no GPU available. 2000ms (single) / 8000ms (ensemble)

Code Deep-Dive

Key algorithms, file structure, and implementation details.

File Structure

src/
├── inference/
│   ├── predict_trinity.py       # Main Trinity predictor (800+ lines)
│   ├── ensemble_expert.py       # 4-model ensemble (300+ lines)
│   ├── micro_intel.py           # Sherlock IMINT (250+ lines)
│   ├── tta_augmentations.py     # Test-time augmentation (150 lines)
│   └── post_processing.py       # Spatial smoothing (100 lines)
├── models/
│   └── geolocator.py            # ViT architecture (400 lines)
├── training/
│   └── train_v4_geoloss.py      # Training script (500 lines)
├── utils/
│   ├── geo_grid.py              # Geographic grid (200 lines)
│   └── istanbul_zones.py        # District mapping (500 lines)
├── app/
│   └── dashboard.py             # Streamlit UI (600 lines)
└── config.py                    # Configuration (180 lines)

Key Files

  • predict_trinity.py: Main entry point
  • ensemble_expert.py: Ensemble logic
  • micro_intel.py: IMINT analysis
  • config.py: All configurations

Lines of Code

  • Total Python: ~4,500 lines
  • Inference logic: ~1,800 lines
  • Training code: ~800 lines
  • Utilities: ~1,200 lines

Key Algorithms

TrinityPredictor.predict()

Main prediction method implementing the 4-layer cascade.

def predict(self, image_path: str) -> dict:
    """Run image through Trinity cascade."""

    # Layer 1: CLIP Guard (landmark detection)
    clip_result = self.clip_guard.analyze(image_path)
    if clip_result["confidence"] > 0.90:
        return {
            "district": clip_result["district"],
            "coordinates": clip_result["coordinates"],
            "confidence": clip_result["confidence"],
            "source": "CLIP Guard"
        }

    # Layer 2: OCR Notary (text extraction)
    ocr_result = self.ocr_notary.extract_text(image_path)
    if ocr_result["confidence"] > 0.70:
        district = self.match_district_keywords(ocr_result["text"])
        if district:
            return {
                "district": district,
                "confidence": ocr_result["confidence"],
                "matched_text": ocr_result["text"],
                "source": "OCR Notary"
            }

    # Layer 3: Ensemble Expert (visual prediction)
    if self.use_ensemble:
        ensemble_result = self.ensemble_expert.predict(image_path)
    else:
        ensemble_result = self.geovit_model.predict(image_path)

    if ensemble_result["confidence"] > 0.60:
        return {
            "district": ensemble_result["district"],
            "coordinates": ensemble_result["coordinates"],
            "confidence": ensemble_result["confidence"],
            "source": "Ensemble Expert" if self.use_ensemble else "GeoViT"
        }

    # Layer 4: Oracle (LLM reasoning with IMINT)
    imint_context = analyze_micro_clues(image_path)  # Sherlock IMINT
    oracle_result = self.oracle.invoke(
        image_path,
        context=imint_context,
        prior_prediction=ensemble_result
    )

    return {
        "district": oracle_result["district"],
        "confidence": oracle_result["confidence"],
        "reasoning": oracle_result["reasoning"],
        "imint_clues": imint_context,
        "source": "Oracle + IMINT"
    }

Sequential Ensemble Loading

Memory-efficient model loading for M3 16GB.

class SequentialEnsemblePredictor:
    def predict(self, image_tensor):
        """Sequential loading to fit M3 16GB memory."""

        model_configs = [
            {"name": "geovit_v4", "path": "models/geovit_v4.pth", "weight": 0.40},
            {"name": "deit_base", "path": "timm/deit_base", "weight": 0.25},
            {"name": "convnext", "path": "timm/convnext_tiny", "weight": 0.20},
            {"name": "efficientnet", "path": "timm/efficientnetv2_s", "weight": 0.15}
        ]

        accumulated_probs = torch.zeros(2000).to("mps")  # 2000 grid cells

        for config in model_configs:
            # Load model to MPS device
            model = self.load_model(config["name"], config["path"])
            model = model.to("mps")
            model.eval()

            # Inference
            with torch.no_grad():
                logits = model(image_tensor.to("mps"))
                probs = torch.softmax(logits, dim=-1)

            # Weighted accumulation
            accumulated_probs += config["weight"] * probs.squeeze()

            # Critical: Immediate cleanup
            del model
            torch.mps.empty_cache()
            gc.collect()  # Force garbage collection

        # Final prediction
        cell_id = accumulated_probs.argmax().item()
        confidence = accumulated_probs[cell_id].item()

        # Convert cell to coordinates
        lat, lon = self.grid.cell_to_coords(cell_id)

        return {
            "cell_id": cell_id,
            "coordinates": (lat, lon),
            "confidence": confidence,
            "probabilities": accumulated_probs
        }

Geographic Loss Function

Distance-aware loss that penalizes far predictions more than nearby ones.

def geo_loss(predictions, true_cell_id, true_coords, alpha=0.3):
    """
    Combine classification loss with geographic distance penalty.

    Args:
        predictions: Model logits (batch_size, num_cells)
        true_cell_id: Ground truth cell ID (batch_size,)
        true_coords: Ground truth (lat, lon) (batch_size, 2)
        alpha: Weight for distance penalty (0-1)

    Returns:
        Combined loss scalar
    """
    # Standard classification loss
    ce_loss = F.cross_entropy(predictions, true_cell_id)

    # Get predicted cell
    pred_cell_id = predictions.argmax(dim=-1)
    pred_coords = grid.cell_to_coords_batch(pred_cell_id)

    # Haversine distance in kilometers
    distances = haversine_distance(pred_coords, true_coords)

    # Distance penalty (normalized by city diameter ~50km)
    distance_penalty = (distances / 50.0).mean()

    # Combined loss
    total_loss = (1 - alpha) * ce_loss + alpha * distance_penalty

    return total_loss

# Example usage in training loop:
# loss = geo_loss(model(images), true_cells, true_coords, alpha=0.3)

Configuration System

All configuration in src/config.py with environment variable overrides.

Environment Variables

# Required API keys (.env file)
GOOGLE_API_KEY=your_gemini_key       # For Gemini Oracle
MAPILLARY_TOKEN=your_mapillary_key   # For data collection
HUGGINGFACE_TOKEN=your_hf_key        # For model uploads

# Feature flags
USE_ENSEMBLE=true                    # Enable ensemble mode
ENABLE_TTA=false                     # Toggle test-time augmentation
ENABLE_IMINT=true                    # Enable Sherlock IMINT

Key Thresholds

Parameter Value Description
CLIP_LANDMARK_THRESHOLD 0.90 Minimum for landmark detection
OCR_CONFIDENCE_THRESHOLD 0.70 Minimum for text acceptance
ENSEMBLE_CONFIDENCE_THRESHOLD 0.60 Triggers Oracle if below
TEMPERATURE_SCALING 1.4 Confidence calibration
SPATIAL_SMOOTHING_ALPHA 0.15 Neighbor blending factor
COORDINATE_REFINEMENT_TOP_K 10 Top cells for centroid

Future Roadmap

Planned improvements and research directions.

Short-term (Next 3 months)

Initiative Description Benefit
Dynamic Threshold Learning Per-district thresholds instead of global 60% Better calibration for each district
Parallel Layer Execution Run CLIP, OCR, and ensemble in parallel Reduce latency by 30-40%
Expanded Landmark Catalog Add 20+ more Istanbul landmarks to CLIP Higher CLIP Guard hit rate
Mobile App Integration Deploy as mobile app with camera input Real-world user testing

Medium-term (6-12 months)

Multi-City Support

Expand to Ankara, Izmir, and other Turkish cities.

  • Challenge: Need new training data and district mappings
  • Approach: Transfer learning from Istanbul model

Real-Time Video Inference

Process video streams frame-by-frame.

  • Challenge: Latency must be <100ms per frame
  • Approach: Model quantization + temporal smoothing

Confidence Calibration Improvements

Better uncertainty estimates using Bayesian methods.

  • Challenge: Current softmax probabilities overconfident
  • Approach: Monte Carlo dropout or ensemble disagreement

Weather/Time-of-Day Layer

Add metadata layer for temporal/weather context.

  • Challenge: Extract time-of-day from shadows, weather from sky
  • Approach: Separate vision model for metadata extraction

Long-term (Research)

Self-Supervised Learning

Train on unlabeled street-view data using contrastive learning.

Potential impact: Scale beyond labeled dataset limitations

Few-Shot City Adaptation

Adapt to new cities with <1000 labeled images.

Potential impact: Rapid deployment to any city worldwide

Multimodal Fusion

Incorporate audio (traffic noise) and depth (LiDAR).

Potential impact: Richer context for geolocation

Edge Deployment

Run entire pipeline on mobile GPU (Apple A-series, Snapdragon).

Potential impact: Offline geolocation without cloud dependency

Open Research Questions

  • Can we reduce ensemble to 2 models (e.g., ViT + CNN) without accuracy loss?
  • How to better prompt-engineer IMINT for structured micro-clue extraction?
  • Can transfer learning generalize to other geolocation tasks (e.g., indoor localization)?
  • What is the theoretical upper bound on visual geolocation accuracy in urban environments?
  • How to handle adversarial cases (intentionally misleading street signs)?
  • Can we predict uncertainty better to know when Oracle will help vs hurt?