职位名称:高级网络爬虫与数据获取工程师 作为 Kaito 的高级网络爬虫与数据获取工程师,您将: - 与联合创始团队紧密合作,定义优先级并制定信息获取路线图 - 设计和实施大规模爬虫系统(100+个爬虫)的架构建设 - 设计、实施并维护数据获取基础设施的各个组件(构建新的爬虫、维护现有爬虫、数据清理器和加载器) - 通过利用或开发统计和机器学习方法来解决大型网络和数据基础设施问题,构建务实的、可扩展的和统计严格的解决方案 - 有效地向研究、工程团队和业务受众推荐技术解决方案 Position: Kaito’s Senior Web Scraping & Data Acquisition Engineer As Kaito’s Senior Web Scraping & Data Acquisition Engineer, you will - Work closely with co-founding team to define priorities and develop information sourcing roadmaps - Lead the effort to design and implement the architecture of a large-scale crawling system (100+ crawlers) - Design, implement, and maintain various components of data acquisition infrastructure (building new crawlers, maintaining existing crawlers, data cleaners & loaders) - Build pragmatic, scalable, and statistically rigorous solutions to large-scale web and data infrastructure problems by leveraging or developing statistical and machine learning methodologies - Effectively advocate technical solutions to research, engineering teams and business audiences
- 具有相关领域的学士学位(例如计算机科学、工程、数学、统计学、运筹学或其他相关领域) - 拥有3年以上使用Python进行数据整理和清洗的经验 - 具备运行、监控和维护爬虫流水线各个方面的专业知识(端到端构建和维护100+个爬虫,避免反爬技术,数据清理和管道化);建议熟悉爬虫库和监控工具(如 BeautifulSoup, Xpaths, Selenium, Puppeteer, Splash) - 拥有从多个不同来源提取数据的经验,包括HTML、XML、REST、GraphQL、PDF和电子表格 - 具备使用技术保护网络爬虫免受机器人检测、网站封禁、IP泄漏、浏览器崩溃、CAPTCHA和代理失败的经验 - 掌握面向对象编程(OOP)、SQL和Django ORM基础知识 加分项: - 具有微服务架构经验 - 具有云环境(例如AWS)经验 - 具有使用容器化工具(如Docker)和编排(如Kubernetes)的经验 - 具有数据仓库维护经验,例如使用AWS Glue Required Qualifications - Bachelor’s degree in quantitative field (e.g. Computer Science, Engineering, Mathematics, Statistics, Operations Research or other related field) - 3+ years of experience with Python for data wrangling and cleaning - Expertise in running, monitoring and maintaining all aspects of a scraping pipeline end to end (building and maintaining 100+ spiders, avoiding bot prevention techniques, data cleaning and pipelining); familiarity with scraping libraries and monitoring tools highly recommended (BeautifulSoup, Xpaths, Selenium, Puppeteer, Splash) - Experience in extracting data from multiple disparate sources including HTML, XML, REST, GraphQL, PDF, and spreadsheets - Experience in using techniques to protect web scrapers against bot detection, site ban, IP leak, browser crash, CAPTCHA and proxy failure - OOP, SQL and Django ORM basics Preferred Qualifications - Experience with micro-services architecture - Experience with cloud environments such as AWS - Experience with containerization tools such as Docker, and orchestration such as kubernetes - Experience with data warehouse maintenance such as AWS Glue 福利待遇: - 全球医保 - 一年两次出国团建 - 弹性工作不打卡 - 有竞争力的薪资 - Global healthcare insurance - Semi-annual international team event - Elastic office hours - Competitive compensation structure