smaple.tr
niranar.com

Data Normalization on niranar.com: Multiple Sources, One DB

Mehmet Kurtipek
January 12, 2026
12 min read
niranar.com
data normalization
case study
Turkey

When a user searches niranar.com for "Bodrum, July 15-22, 4 guests, villa with pool," the results appear within seconds: consistent property names, EUR-denominated prices, precise locations — Yalıkavak, Türkbükü, Göltürkbükü — loaded photos, accurate availability calendars, correct capacity figures. Everything is coherent. Everything speaks the same language.

The order visible on screen is the product of invisible engineering. The platform's 300,000+ active listings are aggregated from external sources, each of which arrived in a different format, with a different vocabulary, a different geographic encoding convention, and a different pricing currency. A clean search result is the output of a normalization process the user never sees. Understanding what that process must solve reveals why building an aggregation platform at this scale is not a data entry project — it is a genuine data engineering challenge.

Why Multi-Source Vacation Rental Data Is Hard

Platforms like Airbnb and Booking.com solve data consistency through centralized control: every host enters data through the same form, validated by the same rules, stored in the same schema. The platform owns the data model. Consistency is enforced at the point of entry.

Aggregating scraped data from external sources works in the opposite direction. The data already exists — in dozens of formats, across dozens of platforms and local agency websites, none of which agreed on a common schema before publishing. One source labels a property "Villa." Another calls it "Summer House," "Independent Rental," or simply "Vacation Property." One provides GPS coordinates. Another provides only a neighborhood name. One prices in EUR; another in Turkish lira; a third in US dollars. One lists capacity in guests; another in bedrooms, leaving the math to the reader.

None of these differences are errors in the source data. Each platform is optimized for its own audience. The problem arises when you want to combine them: left unresolved, 300,000 listings become a pile of disconnected records rather than a coherent, searchable catalog. Filters become unreliable. Maps become inconsistent. Price comparison becomes meaningless.

Every platform that aggregates multi-source data must answer the same set of questions — challenges we also encountered when building kotireitti.fi's multi-source aggregation for the Finnish rental market. How do you unify property type classification when sources use different vocabulary? How do you handle properties with no coordinates, wrong coordinates, or vague location descriptions? How do you identify when the same physical property is listed by multiple sources? How do you reconcile prices in different currencies and with different seasonal structures? How do you manage availability calendars from sources that update at different frequencies? These are simultaneously technical problems and product decisions — every answer has implications for what users see.

The Mediterranean Turkey Context

The challenge is compounded by Turkey's vacation rental market structure, which is considerably more fragmented than markets with mature listing platform ecosystems. Mediterranean Turkey — Antalya, Bodrum, Fethiye, Alanya, Kuşadası, Sapanca — is a high-demand destination for European tourists, particularly from Germany, the UK, Scandinavia, and the Netherlands. These visitors search in a market where availability is spread across independent local agents, small property management companies, family-owned villas listed on obscure regional sites, and a thin layer of global platform coverage that doesn't fully reflect what's actually available on the ground.

This fragmentation is the reason a platform like niranar.com can exist: the gap between what international tourists want to find and what globally standardized platforms actually show them is real. But filling that gap requires solving the aggregation problem that the fragmentation creates. The more heterogeneous the sources, the heavier the normalization burden.

Property Classification: How Villa and Apartment Categories Emerge

niranar.com presents two primary property categories: Villalar (villas) and Daireler (apartments). These categories function as filters in the search interface — when a user selects "Villas," they are communicating a clear intent: private, standalone, likely with outdoor space, probably with a pool.

That clarity doesn't flow automatically from source data. Across different platforms, a property might be labeled "Villa," "Luxury Villa," "Private Villa," "Independent Holiday Home," "Cottage," "Chalet," or a dozen other variants. The apartment category faces similar vocabulary diversity: "Daire," "Rezidans," "Apart," "Stüdyo," "Flat" — any of these might correspond to what niranar.com calls an apartment, depending on context.

Getting this classification right determines whether the platform's core filtering function works. A user searching for villas who receives apartment results loses confidence in the search immediately. The inverse is equally damaging. Classifying 300,000+ listings into two consistent categories requires a mechanism that works regardless of which source used which terminology — one that can resolve ambiguous cases without exposing users to the underlying inconsistency.

When Label Matching Is Not Enough

The vocabulary challenge is deeper in Turkish than in English because certain property terms are genuinely ambiguous. "Yazlık" (literally "summer property") can refer to both villas and apartments depending on the speaker. "Apart" might indicate a specific apartment type or simply be part of a building name. Some listings skip the property type label entirely, using only structural characteristics — number of rooms, pool presence, standalone versus shared building — from which the type must be inferred.

This means classification cannot rely purely on lexical matching. It must interpret context: structural features, amenity descriptions, and the combination of signals that together indicate what kind of property is actually being described. At scale, a classification error isn't a single misplaced listing — it's a systematic error that affects every listing from a particular source or format.

Geographic Consistency: Mapping Mediterranean Turkey

The geographic scope of niranar.com — Antalya, Bodrum, Fethiye, Alanya, Kuşadası, Sapanca — covers some of Turkey's most visited tourist destinations. These are well-known regions to international travelers, but the way they are represented in source data is anything but consistent.

"Bodrum, Yalıkavak" might appear in one source as "Bodrum — Yalıkavak," in another simply as "Muğla," and in a third as "Yalıkavak, 48770." Antalya presents an even more complex case: Lara, Belek, Kemer, and Alanya each function as independent destinations in some sources and as sub-regions of "Antalya" in others. The villages around Fethiye — Ölüdeniz, Hisarönü, Ovacık — appear as standalone destinations in some catalogs and as Fethiye sub-locations in others.

For the map-based browsing experience that Leaflet enables, this geographic inconsistency becomes immediately visible. A property with wrong coordinates appears in the sea or on a mountain. A property with no coordinates disappears from the map entirely while remaining visible in the list view. A property described only as "Bodrum region" gets a center-of-city pin that tells a traveler nothing about where they would actually be staying.

The Leaflet-powered map on niranar.com works consistently across 300,000+ listings. That consistency is the evidence: geographic normalization was solved well enough that the map is a useful navigation tool rather than an unreliable one.

Price Normalization: EUR as the Unifying Currency

All listings on niranar.com are priced in EUR. This provides users with a single, comparable unit: a traveler planning a Bodrum vacation can evaluate a villa alongside an Antalya apartment on the same price basis without performing currency conversion.

Achieving this uniformity requires resolving the currency heterogeneity in source data. Turkish property markets commonly use Turkish lira. International platforms may list in EUR or USD. Some sources may mix currencies across listing types or seasons. Bringing 300,000+ listings to a consistent EUR price involves not just conversion but decisions about which exchange rate applies, how often it is updated, and how to handle historical price data that was originally denominated differently.

Vacation rental pricing adds further complexity. Prices aren't static: high-season rates, low-season rates, weekend premiums, minimum stay requirements, early booking discounts, and last-minute offers all affect the actual price a guest pays. For price comparison to be meaningful across hundreds of thousands of listings from different sources, these parameters need to be represented in a source-agnostic way. A "€50/night" listing that requires a seven-night minimum is a different product from a "€50/night" listing available for a single night — and the filter needs to work correctly in both cases.

niranar.com's price filter — ranging from an accessible tier at approximately €50 and below to a premium segment — demonstrates that the underlying price data is comparable enough to support meaningful segmentation. The filter works because the data is consistent.

Scale: What 300,000+ Means for Data Quality

Normalization problems don't scale linearly. A rule set that handles 100 listings adequately begins to show gaps at 1,000, and requires a substantially different engineering approach at 300,000.

At small scale, human review catches errors. An editor can browse the catalog, notice a misclassified property, and correct it. At 300,000 listings, that approach is unavailable. A flawed classification rule doesn't affect one listing — it affects every listing from a source that uses the vocabulary the rule handles incorrectly. A coordinate resolution error doesn't remove one pin from the map — it removes a region. A currency conversion error doesn't corrupt one price — it corrupts every price in the affected batch.

This means normalization at scale requires rules that handle not just the expected cases but edge cases, ambiguous cases, and genuinely malformed inputs — because at 300,000 records, every possible variant of a problem is represented somewhere in the dataset. The rules also need to fail gracefully: a property with an unresolvable coordinate should not crash the catalog — it should degrade to a reasonable fallback (list view only, no map pin) without contaminating the rest of the data.

Duplicate Detection at Scale

A related problem unique to multi-source aggregation is deduplication: the same physical property may be listed on multiple sources, each with slight variations in name, photos, pricing, or location data. Presenting duplicates to users inflates the apparent catalog size and degrades the browsing experience. Detecting them requires looking past surface-level differences to underlying property identity — which at 300,000 listings is a non-trivial matching problem.

niranar.com's specific deduplication approach is not publicly documented. But the claim of 300,000+ active listings implies that the catalog represents distinct properties, not a count of source records — which in turn implies that the deduplication problem was handled.

Availability and Calendar Normalization

Date-based search is central to the niranar.com experience — users can select check-in and check-out dates, and the platform returns properties that are available for that period. niranar.com implements a custom availability calendar, rather than relying on a third-party calendar service. This keeps the availability model within the platform's control rather than dependent on an external service.

For aggregated listings, availability normalization is a distinct and demanding challenge. Source properties may manage availability in different ways: iCal feeds updated intermittently, API-based availability that reflects live bookings, or manually updated calendars that may lag by days. Different time zone handling across sources can introduce subtle errors. A property that is fully booked for a specific period, but whose availability data is stale, will appear in search results it should not appear in — and users who click through and find the property unavailable will lose trust in the platform.

Availability normalization is less visible than property classification or price normalization, but arguably more operationally critical. A wrong property type classification is a quality problem. A wrong availability result is a direct user experience failure that damages the platform's reliability as a planning tool.

The Evidence: What niranar.com's Output Demonstrates

The internal mechanics of niranar.com's normalization process are not publicly documented, and there is no reason they should be. The specific tools, algorithms, or rule sets applied are operational details that belong to the platform, not its case studies.

What is publicly observable is the output.

The platform presents 300,000+ listings consistently classified as villas or apartments. Properties across Antalya, Bodrum, Fethiye, Alanya, Kuşadası, and Sapanca are geographically positioned correctly enough that Leaflet-powered map navigation is a functional browsing tool. All prices are in EUR and are comparable enough to support meaningful price tier filtering. Calendar-based availability filtering works. Multi-parameter search — location plus date range plus guest count plus amenity filter plus price tier — returns coherent, usable results.

Each of these output characteristics represents a solved normalization problem. The catalog is not a pile of source records. It is a coherent, searchable data structure — which is what normalization produces when done correctly.

What This Signals to Engineering Teams

Multi-source data normalization is the invisible infrastructure of aggregation platforms. Platforms that get it right deliver a transparent user experience: search works, results are meaningful, comparisons are valid, maps are trustworthy. Platforms that don't get it right expose users to the underlying mess — inconsistent categories, wrong map pins, incomparable prices, and availability results that can't be trusted.

niranar.com was built with this distinction in mind. The evidence is in the output: a 300,000-listing catalog that behaves like a single, coherent database rather than a collection of mismatched source files.

For technology teams considering marketplace platforms, aggregation products, or any system that must reconcile data from multiple heterogeneous sources, niranar.com provides a concrete reference point: normalization is not a post-launch cleanup task. It is an architectural decision that must be made at the beginning, because the quality of every downstream feature — search, filtering, maps, pricing, availability — depends on the quality of the data it runs on.

Smart Maple built this data infrastructure as a solo developer using AI-assisted development with Claude Code. That context matters: it demonstrates that getting normalization right at scale doesn't require a large data engineering team — it requires clear product priorities and deliberate architectural choices from day one.

Further Reading

This article is part of a case study series examining how niranar.com works beneath the surface. For the search architecture built on top of this data foundation, see How Search Works on niranar.com. For the technical stack decisions that enable the platform, see niranar.com's Technology Stack. For a product overview, see niranar.com: A Vacation Rental Marketplace for Mediterranean Turkey.

Related Articles

August 11, 2026

MLOps Guide: Taking Machine Learning Models to Production [2026]

87% of machine learning models built by data science teams never reach production. The models work — they pass cross-validation, they score well on holdout sets, they demonstrate genuine predictive value. The problem is not the modeling. The problem is everything that happens between a notebook experiment and a reliable, monitored, production system. MLOps is the discipline that closes that gap. This guide covers the full MLOps stack: maturity levels, tooling choices (MLflow, DVC, Kubeflow

Read More
August 10, 2026

LLM Fine-Tuning Guide: Custom Model Training with LoRA and QLoRA [2026]

General-purpose LLMs are impressive. They can write code, summarize documents, answer questions, and translate between languages with reasonable accuracy. But "reasonable" is not good enough when your application requires consistent output format, domain-specific terminology, a particular tone, or behavior that the base model was never trained to exhibit. That gap is where fine-tuning matters. Fine-tuning updates a model's weights on your specific data, changing how the model behaves — not

Read More
August 9, 2026

Computer Vision Applications: Object Detection, OCR, and Industrial AI [2026]

Computer vision has moved well past the research phase. The models are trained, the frameworks are mature, the hardware is accessible, and the use cases are generating measurable returns. What was a specialized capability requiring deep expertise in 2018 is now deployable infrastructure — if you know which component to reach for and where the real complexity lives. This guide covers computer vision applications across industrial, medical, logistics, and document processing domains. It expl

Read More