Full-Time

Hadoop Admin

Deegit

Deegit

No salary listed

Chicago, IL, USA

In Person

Category
DevOps & Infrastructure (2)
,
Required Skills
Power BI
Microsoft Azure
SAS
R
LDAP
Apache Spark
SQL
Apache Kafka
Tableau
AWS
Hadoop
Yarn
Requirements
  • Knowledgeable in SQL server database.
  • Aggressive self-starter who can operate without infrastructure support.
  • IoT experience.
  • Experience with at least one popular Hadoop distribution — Cloudera, HortonWorks, or MapR.
  • Installation and configuration of Hadoop projects.
Responsibilities
  • Install, configure, and manage Hadoop ecosystem components and related projects such as Hive, Pig, HBase, Spark, etc.
  • Configure and manage YARN and Oozie to achieve optimal resource utilization and end-to-end workflow processing.
  • Deploy and configure graph and document databases such as Neo4J and other emerging big data repositories.
  • Configure data ingestion pipelines using Sqoop, Kafka, and Storm; set up historical and incremental data ingestion processes.
  • Manage file systems including Blob storage and Hadoop Distributed File System and perform backup/restore planning and housekeeping; review and monitor Hadoop log files.
  • Monitor and optimize performance; implement performance tuning and monitoring for Hadoop components and MapReduce jobs.
  • Set up and manage cloud platforms (Azure, AWS, BigTable); deploy big data solutions in cloud environments; perform cluster install and configuration; estimate capacity and configure for elasticity; automate resource allocation and deallocation via scripts.
  • Establish big data security including data at rest and in transit; control access from backend and frontend applications; implement row-level and role-based security; integrate with LDAP, Active Directory, or Okta.
  • Configure front-end access interfaces enabling Cognos, Tableau, Power BI, R, and SAS to access Hadoop big data.
  • Schedule and manage runtime environment using YARN and Oozie for optimal resource utilization and end-to-end workflow processing.
  • Deploy and configure Neo4J (Graph DB) and Document DB and other emerging big data repositories.
  • Engage in big data modeling.
Desired Qualifications
  • Experience with cloud platforms to manage and administer Azure, AWS, BigTable, etc.
  • Deploying Big Data solutions in Cloud.
  • Cluster installation and configuration.
  • Capacity estimation and configuration for elasticity.
  • Script automation for allocation and deallocation of resources.
  • Experience with at least one of Sqoop, Kafka, Storm; knowledge of ADF, Informatica Cloud is a plus.
  • Setup processes for historical and incremental data ingestion.
  • Experience with Hadoop flavors such as Cloudera, HortonWorks, MapR is desirable.
  • Performance optimization for MapReduce jobs and other components.
  • Monitoring and tuning of performance for Hadoop ecosystems.
  • Installation and configuration of Hadoop projects.

Company Size

N/A

Company Stage

N/A

Total Funding

N/A

Headquarters

N/A

Founded

N/A