Frequently asked questions
Can't find your answer? Write to us, we reply quickly.
What is web scraping?
It is the automatic extraction of information published on websites, turned into structured, usable data: a file, a database or an API. Where a person would copy listings one by one, a program does it continuously, without transcription errors and at whatever frequency you need.
What is the collected data actually used for?
To decide on facts rather than instinct: tracking prices in a market, finding prospects, feeding a comparison site or a dashboard, measuring competitors, or spotting an opportunity the moment it is published.
What types of data can you collect?
Publicly accessible data on the web: e-commerce prices and catalogues, real-estate listings, directories and B2B leads, news, and much more depending on your need.
In what format do you deliver the data?
Your choice: CSV or Excel file, REST API, delivered database, or dashboard. We adapt the format to your tools.
Is scraping legal?
We collect publicly accessible data and keep the load on consulted sites under control. We advise you on the framework that applies to your project; responsibility for the purposes of use remains yours.
What happens if a site changes and breaks the collection?
We bring the extractor back up to date. Depending on your plan, this work is scoped by contract.
Are proxies included?
No. Proxies are not included in the subscription and remain your responsibility.
Can you handle large volumes?
Yes, from a few thousand to several million records, both one-off and continuous.
Do you work remotely?
Yes. We support clients throughout France, and beyond.
How does a project start?
With a scoping of the sources, data and desired format, followed by a clear, tailored quote.
Which sites can you collect from?
Public sites in your sector, whichever they are. We work regularly on leboncoin, SeLoger, Bien'ici, PAP, Amazon, Google Maps, PagesJaunes, Indeed, Airbnb, Booking, Tripadvisor and LinkedIn, but that list is in no way exhaustive.
Can I ask for a site that is not on your list?
Yes. Every extractor is built for the sources you name: the list shown on the site is only a sample of the most frequent requests.
How often is the data refreshed?
From a one-off collection to continuous runs, depending on your need. The frequency is agreed during scoping, based on how fast the source changes and how you use the data.
How long before a first delivery?
It depends on the number of sources, their complexity and the protections in place. That is what scoping is for: you know the timeline before development starts.
How do you handle anti-bot protections?
Rotating proxies, captcha handling, stealth and a controlled request rate. The aim is not to force through but to last over time without overloading the sites we consult.
Do you collect data behind a login?
No. We stay with publicly accessible data, with no authentication and no circumvention of restricted access.
Does the collected data contain personal data?
It depends on the sources. Where a collection may contain some, we discuss it during scoping: the fields concerned can be dropped at extraction time. Under the GDPR you remain responsible for how you use the delivered data.
How is data quality checked?
A quality check runs on the output of every collection: field completeness, format consistency and detection of missing or aberrant values, before delivery.
What does the hosting subscription include?
Hosting for your collection solution, keeping it operational, updates and monitoring. Proxies remain at your expense.
How much does a collection cost?
Hosting starts at €150 per month. On-demand scraping and infrastructure built at your premises are quoted per project, based on sources, volume and frequency.
Can I bring the collection in-house later?
Yes, that is exactly what the on-premise offer is for: we build the complete collection chain in your environment, on site or in your cloud, then hand it over to you.
Who owns the collected data?
It is delivered to you in the agreed format and is yours to use for your business. The precise terms are set out in the contract.
Can I get historical data?
A collection captures what is published at the time it runs. History is therefore built by scheduling regular runs, whose results are kept and timestamped.
What happens after delivery?
Depending on the plan: maintenance, monitoring and improvements so collection stays reliable over time, or a full handover if you take it back in-house.
Didn't find your answer?
Ask your question directly. We usually reply within one working day.
Ask a questionLet's talk about your data project
Tell us what you need (sources, volume, frequency, format) and we'll get back to you quickly.