Introduction
A comprehensive technical journey through GeoViT’s architecture, from project inception to cutting-edge ensemble learning and forensic image intelligence. GeoViT is an R&D project at nadcode.
| Training Images | Istanbul Districts | Peak Accuracy | Inference Time |
|---|---|---|---|
| 66,340 | 39 | 81-83% | <1s |
What You’ll Learn
- How we built a 4-layer hierarchical architecture that combines zero-shot learning, OCR, and deep learning
- The evolution from single-model (75% accuracy) to 4-model ensemble (81-83% accuracy)
- Memory-efficient sequential loading strategy for M3 MacBook Air (16GB)
- Sherlock IMINT: Forensic micro-clue analysis (vehicles, flora, infrastructure, signage)
- Research experiments that worked (and those that didn’t) with honest lessons learned
- Complete training pipeline: data collection, geographic grid system, and geo-loss function
Project Timeline
The evolution of GeoViT through three major development phases.
Phase 1: Foundation
January 22, 2026 · 831a6aa
Initial prototype with Trinity Architecture and geographic grid system
Key Achievements
- Built 4-layer Trinity cascade (CLIP → OCR → Visual → Oracle)
- Implemented 2000-cell geographic grid across Istanbul
- Achieved baseline 75% accuracy with single ViT model
- Integrated CLIP for landmark detection, EasyOCR for Turkish text
- Accuracy
- 75%
- Grid Cells
- 2000
- Inference
- 200ms
Phase 2: Ensemble Enhancement
January 31, 2026 · b1eec3a
4-model ensemble with sequential loading, TTA, and post-processing
Key Achievements
- Sequential ensemble: GeoViT (40%), DeiT (25%), ConvNeXt (20%), EfficientNet (15%)
- Memory-efficient loading for M3 16GB (peak 1.5GB)
- Test-Time Augmentation with 8 street-view optimized transforms
- Spatial smoothing + coordinate refinement post-processing
- Accuracy boost: 75% → 81-83% (+6-8 points)
- Accuracy
- 81-83%
- Models
- 4
- Inference
- 1-3s
Phase 3: Sherlock IMINT
Current · In Progress
Forensic micro-clue analysis for low-confidence predictions
Key Achievements
- IMINT (Imagery Intelligence) extracts 4 micro-clue categories
- Vehicles: Turkish plates, yellow taxis, İETT buses, Metrobüs
- Flora: Mediterranean vegetation, coastal vs inland patterns
- Infrastructure: European vs Asian architecture, streetlight styles
- Signage: District logos, Metro markers, directional signs
- Ensemble voting (3 passes) for robust clue detection
- Accuracy Boost
- +2%
- Clue Types
- 4
- Latency
- +3.5s
The Challenge
Visual geolocation in Istanbul presented four critical challenges that shaped our architectural decisions.
Challenge 1: Ambiguous Visual Features
Generic street scenes lack distinctiveness for accurate geolocation.
- Similar apartment buildings across multiple districts
- Generic street furniture (benches, lampposts) provides weak signals
- Visual-only models plateau at ~75% accuracy on ambiguous scenes
- Need for ensemble diversity to capture different feature types
Example: A residential street with standard apartment buildings could exist in Kadıköy, Şişli, or Üsküdar. Without distinct landmarks, visual features alone cannot differentiate.
Challenge 2: Ignoring Text Signals
Vision models treat text as visual patterns, missing semantic ground truth.
- Street signs contain explicit district names (e.g., “Kadıköy Belediyesi”)
- ViT models see text as pixel patterns, not semantic information
- Turkish characters (ı, ğ, ş) require specialized OCR
- Text signals provide 100% ground truth when readable
Solution: OCR Notary layer extracts Turkish text with EasyOCR and overrides visual predictions when district keywords are found with >70% confidence.
Challenge 3: Memory Constraints
Target hardware (M3 MacBook Air 16GB) cannot hold 4 models simultaneously.
- 4 models × ~330MB each = 1.3GB base memory
- Activation tensors during inference add ~500MB per model
- Parallel loading causes OOM errors on 16GB systems
- Need sequential loading: load → predict → unload → next model
Solution: Sequential loading with MPS cache clearing keeps peak memory at ~1.5GB.
for model in [GeoViT, DeiT, ConvNeXt, EfficientNet]:
load_to_mps() → predict() → del model → torch.mps.empty_cache()
Challenge 4: Confidence Calibration
Low-confidence predictions still make guesses instead of asking for help.
- Softmax probabilities are often overconfident (calibration problem)
- 60% confidence on ambiguous scene is unreliable
- Need intelligent fallback when visual models are uncertain
- When to trust vision vs invoke LLM reasoning?
Solution: Oracle layer (Claude/Gemini) activates when ensemble confidence <60%, enriched with IMINT micro-clue context for better reasoning.
Why Istanbul?
Geographic Diversity
- 39 distinct administrative districts
- European vs Asian sides (Bosphorus divide)
- Coastal, urban, suburban, and industrial zones
- Elevation ranges from sea level to 500m hills
Architectural Diversity
- Historic districts (Fatih, Beyoğlu) with Ottoman architecture
- Modern business districts (Maslak, Levent) with skyscrapers
- Residential areas with mixed building styles
- Industrial zones (Tuzla, Pendik) with distinct features
Data Availability
- Rich Mapillary coverage (66,340 street-view images)
- High-quality geotagged data
- Comprehensive district coverage
- Temporal diversity (seasons, time of day)
Perfect Test Bed
- Complex enough to be challenging
- Manageable scope (single city)
- Turkish language OCR requirements
- Real-world application potential
Dataset & Training
How we built and processed 66,340 street-view images into a production-ready geolocation dataset.
Data Sources
Mapillary API
Primary data source for geotagged street-view imagery across Istanbul.
- 66,340 street-view images downloaded via API
- Each image tagged with precise latitude/longitude coordinates
- Coverage across all 39 Istanbul districts
- Temporal diversity: different seasons, times of day, weather conditions
Geographic Grid System
Instead of regression (predicting continuous lat/lon), we use classification with a 2000-cell grid. This approach provides better generalization and more stable training.
Grid Specifications
- 2000 cells across Istanbul
- Cell size: ~150m × 150m
- Bounding box:
(28.4°W, 40.75°S, 29.70°E, 41.30°N) - Covers all 39 districts comprehensively
Why Classification?
- More stable training than regression
- Better generalization on test set
- Easier to interpret model confidence
- Allows spatial smoothing post-processing
Conversion Process
Raw Image (lat=41.0082, lon=28.9784)
↓ Geocoding
Grid Cell ID: 1247 (Beşiktaş district)
↓ Training Target
One-hot vector [0, 0, ..., 1, ..., 0] (2000 dims)
↓ Inference
Predicted Cell: 1247 → Convert to (41.008, 28.978)
Data Processing Pipeline
- Raw Images (Mapillary): Download via API with geotags
- Geocoding: Convert lat/lon → grid cell ID using bounding box math
- Train/Val/Test Split: 80/10/10 stratified by district for balanced representation
- Augmentation: Flip, rotate, color jitter, random crop during training
- ViT Preprocessing: Resize to 224×224, normalize with ImageNet stats
Training Process
Training Configuration
- Loss function: Geo-loss (classification + distance penalty)
- Optimizer: AdamW with cosine schedule
- Batch size: 32 (M3 memory constraint)
- Epochs: 15-20 until convergence
- Hardware: M3 MacBook Air (MPS backend)
Geo-Loss Function
Combines classification loss with geographic distance penalty to prioritize nearby cells over distant ones.
loss = cross_entropy(pred, true_cell)
+ α * distance_penalty(pred, true_coords)
# Prefer wrong nearby cells
# over wrong distant cells
Training Command
python3 -m src.training.train_v4_geoloss \
--data_dir data/processed/istanbul \
--epochs 15 \
--batch_size 32 \
--lr 1e-4 \
--model_name geovit_v4
Dataset Statistics
| Metric | Value | Notes |
|---|---|---|
| Total Images | 66,340 | 80/10/10 train/val/test split, stratified by district |
| Districts Covered | 39 | Balanced representation across all districts |
| Grid Cells | 2000 | ~150m × 150m cell resolution |
Trinity Architecture
A 4-layer hierarchical cascade where each layer can override the next, combining zero-shot learning, OCR, deep learning, and LLM reasoning.
Core Philosophy
The Trinity Architecture uses a hierarchical override cascade where higher layers can short-circuit lower layers based on confidence thresholds. This ensures fast inference on easy cases while falling back to more expensive reasoning when needed.
Why “Trinity”? Originally referred to the core 3 layers (CLIP, OCR, Visual). The Oracle was added later as a fourth fallback layer for low-confidence predictions.
Decision Flow
CLIP Guard
Landmark detected? Threshold: 90%
Return landmark location
Pass to OCR
OCR Notary
District text found? Threshold: 70%
Return text-based location
Pass to Ensemble Expert
Ensemble Expert
Visual confidence high? Threshold: 60%
Return visual prediction
Pass to Oracle
Oracle
LLM reasoning Threshold: Fallback
Return LLM prediction
Cascade Logic
Each layer processes the image in order. If a layer meets its confidence threshold, it returns immediately. Otherwise, the image cascades to the next layer. The Oracle is the final fallback for low-confidence predictions.Layer 1: CLIP Guard
Zero-shot landmark detection for famous Istanbul locations.
How It Works
- Model: OpenAI CLIP (ViT-B/32)
- Method: Text-image similarity scoring
- Threshold: 90% confidence
- Inference: ~0.1s
- Landmarks: 12 famous locations
Examples
- Galata Tower (Beyoğlu)
- Hagia Sophia (Fatih)
- Blue Mosque (Fatih)
- Maiden’s Tower (Üsküdar)
- Bosphorus Bridge (Beşiktaş/Üsküdar)
- Dolmabahçe Palace (Beşiktaş)
Code Snippet
# CLIP Guard inference loop
for landmark_name, config in ISTANBUL_LANDMARKS.items():
prompts = config["prompts"] # e.g., ["a photo of Galata Tower"]
similarity = clip_model(image, prompts)
if similarity > 0.90: # High confidence threshold
return {
"district": config["district"],
"coordinates": config["coordinates"],
"confidence": similarity,
"source": "CLIP Guard"
}
Limitation: Only works for famous landmarks with distinctive visual features. Generic street scenes pass through to next layer.
Layer 2: OCR Notary
Extract Turkish text from signs and match district keywords.
How It Works
- Model: EasyOCR (Turkish + English)
- Method: Text extraction + keyword matching
- Threshold: 70% OCR confidence
- Inference: ~0.3s
- Keywords: District names, Metro stations
Examples
- “Kadıköy Belediyesi” → Kadıköy
- “Beşiktaş Metro” → Beşiktaş
- “Şişli AVM” → Şişli
- “Üsküdar İskele” → Üsküdar
- “Fatih Municipality” → Fatih
Code Snippet
# OCR Notary text extraction
results = ocr_reader.readtext(image_array, paragraph=False)
for (bbox, text, confidence) in results:
if confidence < 0.70: # Skip low-confidence detections
continue
text_clean = text.lower().strip()
# Match against district keywords
for district, keywords in DISTRICT_KEYWORDS.items():
if any(kw in text_clean for kw in keywords):
return {
"district": district,
"confidence": confidence,
"matched_text": text,
"source": "OCR Notary"
}
Limitation: Requires readable text in the image. Blurry, occluded, or non-text images pass through to next layer.
Layer 3: Ensemble Expert
Visual pattern recognition with ensemble learning (optional).
Single Model Mode
- Model: GeoViT v4 (ViT-base)
- Training: Geo-loss on 66,340 images
- Threshold: 60% confidence
- Inference: ~0.2s
- Accuracy: 75%
Ensemble Mode
- Models: GeoViT, DeiT, ConvNeXt, EfficientNet
- Strategy: Sequential loading + weighted voting
- Inference: ~1-3s (with TTA)
- Accuracy: 81-83%
- Memory: Peak 1.5GB
Code Snippet (Ensemble)
# Sequential ensemble loading (memory-efficient)
model_configs = [
{"name": "geovit_v4", "weight": 0.40},
{"name": "deit_base", "weight": 0.25},
{"name": "convnext_tiny", "weight": 0.20},
{"name": "efficientnet_v2_s", "weight": 0.15}
]
accumulated_probs = torch.zeros(2000) # 2000 grid cells
for config in model_configs:
model = load_model(config["name"]).to("mps")
probs = model(image_tensor)
accumulated_probs += config["weight"] * probs
del model # Unload from memory
torch.mps.empty_cache() # Clear MPS cache
final_prediction = accumulated_probs.argmax()
Limitation: Ambiguous scenes with low visual distinctiveness (<60% confidence) trigger the Oracle layer.
Layer 4: Oracle
LLM-based reasoning with IMINT context (low-confidence fallback).
How It Works
- Models: Claude CLI or Gemini 2.0
- Trigger: Ensemble confidence <60%
- Context: IMINT micro-clues (see Sherlock IMINT)
- Inference: ~3.5s
- Use case: Ambiguous residential streets
IMINT Integration
- Extracts vehicles (taxis, buses, plates)
- Analyzes flora (Mediterranean vegetation)
- Detects infrastructure (architecture style)
- Reads signage (district logos, Metro)
- Provides structured context to LLM
Code Snippet
# Oracle invocation with IMINT context
if ensemble_confidence < 0.60:
# Extract micro-clues
imint_context = analyze_micro_clues(image_path)
# Build prompt with context
prompt = f"""
Analyze this Istanbul street-view image.
Ensemble prediction: {ensemble_district} ({ensemble_confidence:.1%})
IMINT Context:
- Vehicles: {imint_context['vehicles']}
- Flora: {imint_context['flora']}
- Infrastructure: {imint_context['infrastructure']}
- Signage: {imint_context['signage']}
Which Istanbul district is this most likely?
"""
oracle_result = llm_client.invoke(prompt, image=image_path)
return parse_oracle_response(oracle_result)
Performance: Oracle adds ~3.5s latency but improves accuracy by ~2% on low-confidence predictions. Only activated for ~15% of images.
Ensemble Evolution
From single-model baseline to 4-model ensemble with sequential loading, TTA, and post-processing.
Why Ensemble?
Single models plateau at 75% accuracy because they learn similar features. Ensemble diversity allows each model to capture different aspects of the scene.
The Problem
- Single GeoViT v4: 75% accuracy ceiling
- High confidence on wrong predictions
- Struggles with ambiguous residential streets
- Similar errors across training runs
The Solution
- 4 diverse models with different architectures
- Weighted soft voting combines strengths
- Ensemble: 81-83% accuracy (+6-8 points)
- More robust confidence calibration
Model Selection
| Model | Weight | Description | Model Size | Key Strength |
|---|---|---|---|---|
| GeoViT v4 | 40% | Anchor model trained with geo-loss. Best single-model performance. | ~330 MB | Fine-tuned on Istanbul, strong on landmarks |
| DeiT-Base | 25% | Distilled Vision Transformer. Efficient and generalizes well. | ~330 MB | Good on generic scenes, robust to variations |
| ConvNeXt-Tiny | 20% | Modernized ConvNet. Strong local pattern recognition. | ~110 MB | Captures architectural details, street patterns |
| EfficientNetV2-S | 15% | Mobile-optimized CNN. Fast and memory-efficient. | ~85 MB | Good on texture patterns, building materials |
Why These 4 Models?
Selected through ablation studies. Each model contributes unique features:
- Architecture diversity: ViT (GeoViT, DeiT) + CNN (ConvNeXt, EfficientNet)
- Scale diversity: Different model sizes capture features at different scales
- Training diversity: Different initialization and augmentation strategies
- Complementary errors: Models fail on different subsets of images
Sequential Loading Strategy
The Memory Challenge
Target hardware: M3 MacBook Air with 16GB unified memory. Loading all 4 models simultaneously causes OOM (Out of Memory) errors.
# Parallel loading (FAILS on M3 16GB)
models = [
GeoViT().to("mps"), # ~330MB + 400MB activations
DeiT().to("mps"), # ~330MB + 400MB activations
ConvNeXt().to("mps"), # ~110MB + 300MB activations
EfficientNet().to("mps") # ~85MB + 250MB activations
]
# Total: ~2.2GB - causes OOM with other processes
# Sequential loading (WORKS on M3 16GB)
for model_config in ensemble_models:
model = load_model(config).to("mps") # ~800MB peak
probs = model(image)
accumulated += config["weight"] * probs
del model
torch.mps.empty_cache() # Release memory
# Peak: ~1.5GB - fits comfortably
| Peak Memory | Speed | Fits M3 16GB |
|---|---|---|
| 1.5GB | 4× slower (vs parallel) | Yes |
Test-Time Augmentation (TTA)
Apply 8 street-view optimized augmentations at inference time and average predictions for improved robustness.
| Augmentation | Purpose |
|---|---|
| Horizontal Flip | Mimic reverse driving direction |
| Rotate ±5° | Camera angle variations |
| Random Crop | Different framing |
| Brightness ±20% | Time of day variations |
| Contrast ±15% | Weather conditions |
| Saturation ±10% | Color calibration |
| Hue ±5° | White balance shifts |
| Gaussian Blur | Motion blur simulation |
TTA Strategy
# Average predictions across augmentations
tta_probs = []
for augmentation in [flip, rotate, crop, brightness, ...]:
aug_image = augmentation(original_image)
probs = ensemble_predict(aug_image)
tta_probs.append(probs)
final_probs = torch.stack(tta_probs).mean(dim=0) # Average
prediction = final_probs.argmax()
- Benefit: +2.5% accuracy improvement on test set
- Cost: 8× inference time (3s total with ensemble)
Post-Processing Pipeline
Spatial Smoothing
Blend probabilities of neighbor cells (α=0.15) to create smoother predictions.
smoothed[cell] = (1-α) * pred[cell] + α * mean(neighbors)
Impact: +0.5% accuracy
Coordinate Refinement
Weighted centroid of top-10 cells instead of single cell center.
coords = Σ(top_k_cells * their_probs) / Σ(top_k_probs)
Impact: +0.3% accuracy
Temperature Scaling
Calibrate confidence scores with T=1.4 (found via grid search).
calibrated_probs = softmax(logits / T)
Impact: Better uncertainty estimates
Performance Comparison
| Configuration | Accuracy | Inference Time | Use Case |
|---|---|---|---|
| Single GeoViT v4 | 75.0% | 200ms | Real-time apps |
| Ensemble (no TTA) | 78.5% | 1000ms | Production (recommended) |
| Ensemble + TTA | 81.0% | 3000ms | Batch processing |
| Ensemble + TTA + PostP | 81.5% | 3200ms | Maximum accuracy |
| + IMINT (low conf) | 83.0% | 6700ms | Critical accuracy |
Sherlock IMINT
Forensic imagery intelligence that extracts micro-clues invisible to ML models.
What is IMINT?
IMINT (Imagery Intelligence) is a military and forensic technique for extracting actionable information from images. We adapt this methodology to visual geolocation by analyzing micro-clues that ML models overlook.
Traditional Approach (Vision Models)
- Learn pixel patterns end-to-end
- Miss semantic details (license plate text)
- No explicit reasoning process
- Low confidence on ambiguous scenes
IMINT Approach (Sherlock)
- Extract structured micro-clues
- Semantic understanding (34 = Istanbul plate)
- Explicit reasoning chain
- Provides context to Oracle LLM
Why Sherlock IMINT?
The Problem
When ensemble confidence is low (<60%), we invoke the Oracle LLM. But generic “analyze this image” prompts waste the LLM’s reasoning capacity on basic visual tasks.
The Solution
Sherlock IMINT pre-processes the image to extract structured micro-clues, then provides this context to the Oracle. This focuses the LLM’s attention on high-level reasoning rather than low-level visual parsing.
Impact
| Accuracy Boost | Latency Added | Images Triggered |
|---|---|---|
| +2% | +3.5s | ~15% |
The 4 Micro-Clue Categories
Key indicators for geolocation, grouped into four categories.
1. Vehicles
- Turkish license plates (34 prefix = Istanbul)
- Yellow taxis (ubiquitous in Istanbul)
- İETT buses (public transport branding)
- Metrobüs (BRT system, specific routes)
- Dolmuş minibuses (district-specific colors)
Analysis prompt:
Look for vehicles, license plates (34 prefix?), taxis (yellow?), buses (İETT branding?), Metrobüs.
2. Flora
- Mediterranean vegetation patterns
- Coastal palm trees (Bosphorus areas)
- Plane trees (common in European side)
- Pine trees (Asian side hills)
- Seasonal indicators (spring blossoms)
Analysis prompt:
Analyze vegetation: Mediterranean species? Palm trees (coastal)? Pine trees (Asian side)? Plane trees (European side)?
3. Infrastructure
- European vs Asian side architecture
- Modern (Maslak/Levent skyscrapers)
- Historic (Fatih/Beyoğlu Ottoman buildings)
- Streetlight styles (vary by district)
- Pavement patterns and materials
- Power line configurations
Analysis prompt:
Examine architecture: Modern vs historic? European vs Asian style? Streetlight design? Pavement type?
4. Signage
- Blue directional signs (Istanbul standard)
- Municipality logos (district branding)
- Metro/tramvay station markers
- AVM (shopping mall) names
- District boundary signs
Analysis prompt:
Read all visible text: District names? Municipality logos? Metro stations? Street signs? Mall names?
Implementation
Sherlock IMINT uses ensemble voting (3 passes with Claude CLI) to extract robust micro-clues, then provides structured context to the Oracle.
def analyze_micro_clues(image_path: str) -> dict:
"""Extract IMINT micro-clues via ensemble voting."""
clue_categories = ["vehicles", "flora", "infrastructure", "signage"]
results = {cat: [] for cat in clue_categories}
# Run 3 independent passes for robustness
for pass_idx in range(3):
for category in clue_categories:
prompt = IMINT_PROMPTS[category] # Category-specific prompt
response = claude_cli.invoke(prompt, image=image_path)
parsed = parse_imint_response(response, category)
results[category].append(parsed)
# Majority voting
consensus = {}
for category in clue_categories:
# Find most common answer across 3 passes
consensus[category] = majority_vote(results[category])
return consensus
# Example output:
# {
# "vehicles": "Yellow taxis, 34 license plates, İETT bus visible",
# "flora": "Mediterranean plane trees, coastal vegetation",
# "infrastructure": "Modern streetlights, European-style pavement",
# "signage": "Beşiktaş Municipality logo, Metro sign"
# }
Ensemble Voting Strategy
- Run analysis 3 times independently
- Majority vote on each clue category
- Filters out hallucinations and noise
- More robust than single-pass
Oracle Integration
- IMINT context provided to Oracle LLM
- Structured JSON format for parsing
- LLM focuses on high-level reasoning
- Improves accuracy on ambiguous cases
Performance Impact
| Metric | Value | Notes |
|---|---|---|
| Accuracy Boost | +2% | On low-confidence predictions where Oracle is invoked |
| Latency Added | +3.5s | 3 Claude CLI passes + parsing overhead |
| Images Triggered | ~15% | Only activated when ensemble confidence <60% |
Example Scenario
Image: Generic residential street with no landmarks.
- Before: Ensemble: Kadıköy (58% confidence) → Oracle invoked → Generic analysis → Wrong: Şişli
- After: IMINT: “Yellow taxi, 34 plate, İETT bus, European architecture” → Oracle with context → Correct: Kadıköy
What Worked & What Didn’t
Honest lessons from research experiments, including failures and pivots.
Experiments That Worked
Geographic Loss Function
- Problem: Standard cross-entropy treats all wrong predictions equally
- Solution: Geo-loss penalizes distant predictions more than nearby ones
- Result: +3% accuracy improvement
loss = cross_entropy(pred, true_cell) + α * distance_penalty(pred_coords, true_coords)
Sequential Ensemble Loading
- Problem: Parallel loading causes OOM on M3 16GB
- Solution: Load → predict → unload → next model
- Result: Memory: 3GB → 1.5GB (fits M3)
for model in models:
load() → predict() → del model → torch.mps.empty_cache()
OCR Override Strategy
- Problem: Vision models ignore text semantics
- Solution: Trust OCR text when confidence >70%, override visual predictions
- Result: +5% accuracy on text-heavy images
if ocr_confidence > 0.70:
return ocr_district # Override visual prediction
Temperature Scaling
- Problem: Softmax probabilities are overconfident
- Solution: Calibrate with T=1.4 (found via grid search)
- Result: Better uncertainty estimates, more reliable confidence scores
calibrated_probs = softmax(logits / 1.4)
Experiments That Failed
Regression-Based Approach
- What we tried: Direct lat/lon prediction instead of classification
- The problem: Poor generalization, high variance in predictions
- Why it failed: Continuous space too large, model struggles to converge
- Switched to: Classification with 2000-cell grid → stable training
Parallel Ensemble Inference
- What we tried: Load all 4 models simultaneously for faster inference
- The problem: OOM (Out of Memory) errors on M3 16GB
- Why it failed: 4 models × 330MB + activations > 2GB memory
- Switched to: Sequential loading → fits memory, 4× slower but works
Random TTA Transforms
- What we tried: Apply generic augmentations (vertical flip, extreme rotation)
- The problem: Degraded accuracy by -1%
- Why it failed: Street-view has specific geometry (no upside-down scenes)
- Switched to: Curated 8 street-view specific transforms
Attention Pooling
- What we tried: Attention-based ensemble aggregation instead of voting
- The problem: Slower inference, no accuracy gain
- Why it failed: Overcomplicates the problem, simple voting works well
- Switched to: Weighted soft voting → simpler and just as good
Lessons Learned
- Simplicity often wins: Weighted soft voting outperforms complex attention mechanisms
- Text signals are gold: OCR override provides 100% ground truth when readable
- Memory constraints drive architecture: Target hardware shapes design decisions
- Domain-specific augmentations matter: Generic TTA can hurt more than help
- Classification > Regression: For geolocation, discrete cells generalize better than continuous coords
- Failed experiments are valuable: Each failure taught us what NOT to do
Performance Optimization
Balancing accuracy, speed, and memory constraints for production deployment.
Memory Optimization
Target hardware: M3 MacBook Air with 16GB unified memory. All optimizations focus on fitting within this constraint.
Techniques
- Sequential loading: One model at a time
- MPS backend: Metal Performance Shaders on M3
- Batch size tuning: 32 optimal for training
- Gradient checkpointing: During training
- Cache clearing:
torch.mps.empty_cache()
Memory Profile
| Configuration | Peak Memory |
|---|---|
| Single Model | ~800MB |
| Ensemble (Sequential) | ~1.5GB |
| Ensemble (Parallel, OOM) | ~3.2GB |
Memory Management Code
# Proper memory management for MPS
import torch
for model_config in ensemble_models:
# Load model to MPS device
model = load_model(config["name"]).to("mps")
# Inference
with torch.no_grad():
probs = model(image_tensor)
# Accumulate weighted probabilities
accumulated_probs += config["weight"] * probs
# Critical: Clean up immediately
del model
torch.mps.empty_cache() # Release Metal memory
# Peak memory stays under 1.5GB
Speed Optimization
Early Exit Strategies
Trinity layers short-circuit when high confidence is achieved, avoiding expensive inference.
- CLIP Guard: ~0.1s (if landmark detected)
- OCR Notary: ~0.3s (if text found)
- Ensemble: ~1-3s (most images)
- Oracle: +3.5s (only ~15% of images)
Caching Strategies
Pre-compute and cache static resources to avoid redundant work.
- CLIP embeddings: Cache for 12 landmarks
- OCR keywords: Pre-compiled regex
- Grid coordinates: Pre-computed lookup table
- Model weights: Load once, reuse
Latency Breakdown (Average Image)
| Task | Time |
|---|---|
| CLIP Guard check | 100ms |
| OCR Notary extraction | 300ms |
| Ensemble inference (4 models) | 1000ms |
| Post-processing | 50ms |
| Coordinate conversion | 10ms |
Accuracy vs Speed Tradeoffs
| Mode | Accuracy | Latency | Use Case | Recommended |
|---|---|---|---|---|
| Single GeoViT | 75% | 200ms | Real-time apps, demos | |
| Ensemble (no TTA) | 78.5% | 1000ms | Production deployment | Best |
| Ensemble + TTA | 81.5% | 3000ms | Batch processing | |
| + IMINT (low conf) | 83% | 6500ms | Critical accuracy needs |
Recommendation: Ensemble (no TTA) provides the best accuracy/latency tradeoff for production. TTA and IMINT can be enabled selectively for high-value predictions.
Device Compatibility
| Device | Status | Description | Inference Time |
|---|---|---|---|
| MPS (M1/M2/M3) | Optimized | Primary target. Metal Performance Shaders on Apple Silicon. | 200ms (single) / 1000ms (ensemble) |
| CUDA (NVIDIA GPU) | Supported | Supported for users with NVIDIA GPUs. | 150ms (single) / 800ms (ensemble) |
| CPU | Slow | Fallback mode when no GPU available. | 2000ms (single) / 8000ms (ensemble) |
Code Deep-Dive
Key algorithms, file structure, and implementation details.
File Structure
src/
├── inference/
│ ├── predict_trinity.py # Main Trinity predictor (800+ lines)
│ ├── ensemble_expert.py # 4-model ensemble (300+ lines)
│ ├── micro_intel.py # Sherlock IMINT (250+ lines)
│ ├── tta_augmentations.py # Test-time augmentation (150 lines)
│ └── post_processing.py # Spatial smoothing (100 lines)
├── models/
│ └── geolocator.py # ViT architecture (400 lines)
├── training/
│ └── train_v4_geoloss.py # Training script (500 lines)
├── utils/
│ ├── geo_grid.py # Geographic grid (200 lines)
│ └── istanbul_zones.py # District mapping (500 lines)
├── app/
│ └── dashboard.py # Streamlit UI (600 lines)
└── config.py # Configuration (180 lines)
Key Files
predict_trinity.py: Main entry pointensemble_expert.py: Ensemble logicmicro_intel.py: IMINT analysisconfig.py: All configurations
Lines of Code
- Total Python: ~4,500 lines
- Inference logic: ~1,800 lines
- Training code: ~800 lines
- Utilities: ~1,200 lines
Key Algorithms
TrinityPredictor.predict()
Main prediction method implementing the 4-layer cascade.
def predict(self, image_path: str) -> dict:
"""Run image through Trinity cascade."""
# Layer 1: CLIP Guard (landmark detection)
clip_result = self.clip_guard.analyze(image_path)
if clip_result["confidence"] > 0.90:
return {
"district": clip_result["district"],
"coordinates": clip_result["coordinates"],
"confidence": clip_result["confidence"],
"source": "CLIP Guard"
}
# Layer 2: OCR Notary (text extraction)
ocr_result = self.ocr_notary.extract_text(image_path)
if ocr_result["confidence"] > 0.70:
district = self.match_district_keywords(ocr_result["text"])
if district:
return {
"district": district,
"confidence": ocr_result["confidence"],
"matched_text": ocr_result["text"],
"source": "OCR Notary"
}
# Layer 3: Ensemble Expert (visual prediction)
if self.use_ensemble:
ensemble_result = self.ensemble_expert.predict(image_path)
else:
ensemble_result = self.geovit_model.predict(image_path)
if ensemble_result["confidence"] > 0.60:
return {
"district": ensemble_result["district"],
"coordinates": ensemble_result["coordinates"],
"confidence": ensemble_result["confidence"],
"source": "Ensemble Expert" if self.use_ensemble else "GeoViT"
}
# Layer 4: Oracle (LLM reasoning with IMINT)
imint_context = analyze_micro_clues(image_path) # Sherlock IMINT
oracle_result = self.oracle.invoke(
image_path,
context=imint_context,
prior_prediction=ensemble_result
)
return {
"district": oracle_result["district"],
"confidence": oracle_result["confidence"],
"reasoning": oracle_result["reasoning"],
"imint_clues": imint_context,
"source": "Oracle + IMINT"
}
Sequential Ensemble Loading
Memory-efficient model loading for M3 16GB.
class SequentialEnsemblePredictor:
def predict(self, image_tensor):
"""Sequential loading to fit M3 16GB memory."""
model_configs = [
{"name": "geovit_v4", "path": "models/geovit_v4.pth", "weight": 0.40},
{"name": "deit_base", "path": "timm/deit_base", "weight": 0.25},
{"name": "convnext", "path": "timm/convnext_tiny", "weight": 0.20},
{"name": "efficientnet", "path": "timm/efficientnetv2_s", "weight": 0.15}
]
accumulated_probs = torch.zeros(2000).to("mps") # 2000 grid cells
for config in model_configs:
# Load model to MPS device
model = self.load_model(config["name"], config["path"])
model = model.to("mps")
model.eval()
# Inference
with torch.no_grad():
logits = model(image_tensor.to("mps"))
probs = torch.softmax(logits, dim=-1)
# Weighted accumulation
accumulated_probs += config["weight"] * probs.squeeze()
# Critical: Immediate cleanup
del model
torch.mps.empty_cache()
gc.collect() # Force garbage collection
# Final prediction
cell_id = accumulated_probs.argmax().item()
confidence = accumulated_probs[cell_id].item()
# Convert cell to coordinates
lat, lon = self.grid.cell_to_coords(cell_id)
return {
"cell_id": cell_id,
"coordinates": (lat, lon),
"confidence": confidence,
"probabilities": accumulated_probs
}
Geographic Loss Function
Distance-aware loss that penalizes far predictions more than nearby ones.
def geo_loss(predictions, true_cell_id, true_coords, alpha=0.3):
"""
Combine classification loss with geographic distance penalty.
Args:
predictions: Model logits (batch_size, num_cells)
true_cell_id: Ground truth cell ID (batch_size,)
true_coords: Ground truth (lat, lon) (batch_size, 2)
alpha: Weight for distance penalty (0-1)
Returns:
Combined loss scalar
"""
# Standard classification loss
ce_loss = F.cross_entropy(predictions, true_cell_id)
# Get predicted cell
pred_cell_id = predictions.argmax(dim=-1)
pred_coords = grid.cell_to_coords_batch(pred_cell_id)
# Haversine distance in kilometers
distances = haversine_distance(pred_coords, true_coords)
# Distance penalty (normalized by city diameter ~50km)
distance_penalty = (distances / 50.0).mean()
# Combined loss
total_loss = (1 - alpha) * ce_loss + alpha * distance_penalty
return total_loss
# Example usage in training loop:
# loss = geo_loss(model(images), true_cells, true_coords, alpha=0.3)
Configuration System
All configuration in src/config.py with environment variable overrides.
Environment Variables
# Required API keys (.env file)
GOOGLE_API_KEY=your_gemini_key # For Gemini Oracle
MAPILLARY_TOKEN=your_mapillary_key # For data collection
HUGGINGFACE_TOKEN=your_hf_key # For model uploads
# Feature flags
USE_ENSEMBLE=true # Enable ensemble mode
ENABLE_TTA=false # Toggle test-time augmentation
ENABLE_IMINT=true # Enable Sherlock IMINT
Key Thresholds
| Parameter | Value | Description |
|---|---|---|
CLIP_LANDMARK_THRESHOLD |
0.90 | Minimum for landmark detection |
OCR_CONFIDENCE_THRESHOLD |
0.70 | Minimum for text acceptance |
ENSEMBLE_CONFIDENCE_THRESHOLD |
0.60 | Triggers Oracle if below |
TEMPERATURE_SCALING |
1.4 | Confidence calibration |
SPATIAL_SMOOTHING_ALPHA |
0.15 | Neighbor blending factor |
COORDINATE_REFINEMENT_TOP_K |
10 | Top cells for centroid |
Future Roadmap
Planned improvements and research directions.
Short-term (Next 3 months)
| Initiative | Description | Benefit |
|---|---|---|
| Dynamic Threshold Learning | Per-district thresholds instead of global 60% | Better calibration for each district |
| Parallel Layer Execution | Run CLIP, OCR, and ensemble in parallel | Reduce latency by 30-40% |
| Expanded Landmark Catalog | Add 20+ more Istanbul landmarks to CLIP | Higher CLIP Guard hit rate |
| Mobile App Integration | Deploy as mobile app with camera input | Real-world user testing |
Medium-term (6-12 months)
Multi-City Support
Expand to Ankara, Izmir, and other Turkish cities.
- Challenge: Need new training data and district mappings
- Approach: Transfer learning from Istanbul model
Real-Time Video Inference
Process video streams frame-by-frame.
- Challenge: Latency must be <100ms per frame
- Approach: Model quantization + temporal smoothing
Confidence Calibration Improvements
Better uncertainty estimates using Bayesian methods.
- Challenge: Current softmax probabilities overconfident
- Approach: Monte Carlo dropout or ensemble disagreement
Weather/Time-of-Day Layer
Add metadata layer for temporal/weather context.
- Challenge: Extract time-of-day from shadows, weather from sky
- Approach: Separate vision model for metadata extraction
Long-term (Research)
Self-Supervised Learning
Train on unlabeled street-view data using contrastive learning.
Potential impact: Scale beyond labeled dataset limitations
Few-Shot City Adaptation
Adapt to new cities with <1000 labeled images.
Potential impact: Rapid deployment to any city worldwide
Multimodal Fusion
Incorporate audio (traffic noise) and depth (LiDAR).
Potential impact: Richer context for geolocation
Edge Deployment
Run entire pipeline on mobile GPU (Apple A-series, Snapdragon).
Potential impact: Offline geolocation without cloud dependency
Open Research Questions
- Can we reduce ensemble to 2 models (e.g., ViT + CNN) without accuracy loss?
- How to better prompt-engineer IMINT for structured micro-clue extraction?
- Can transfer learning generalize to other geolocation tasks (e.g., indoor localization)?
- What is the theoretical upper bound on visual geolocation accuracy in urban environments?
- How to handle adversarial cases (intentionally misleading street signs)?
- Can we predict uncertainty better to know when Oracle will help vs hurt?