Business Intelligence v1: data lake
- Type
- Project
TrueStory's first attempt at a company-wide data lake, and my role was foundational: research, not execution against an existing plan, because there wasn't one yet.
I did the initial investigation into what building a data lake for all of TrueStory's data would actually take, documented findings and technology-selection recommendations for the team to act on, and then went further than a slide deck, building proof-of-concept solutions to validate the approach before anyone committed real engineering time to it.
The stack under consideration was AWS-native:
- Athena for querying
- Glue and Glue DataBrew for transformation
- S3 as the lake itself
- Lake Formation for access governance
- EventBridge for event routing
- Secrets Manager for credentials
Alongside the technology research, I worked directly on data denormalization, transformation, and enrichment, the unglamorous work of actually shaping raw data into something a data lake could usefully hold. It's the project in my TrueStory stretch that looked least like typical backend work and most like the AI/data-engineering direction my career eventually turned toward.
Built with
- C#
- MySQL
- Microsoft SQL Server
- AWS Glue / Athena
- AWS EventBridge
- Amazon S3
- AWS IAM
- AWS Lake Formation
- AWS Glue DataBrew
- AWS Secrets Manager
- SonarQube
- Bitbucket (Repositories, Pipelines)
- Slack
- Rancher
- Traefik
- Spot by NetApp
- LogStash / Kibana (ELK)
- JIRA
- Shape Up methodology
- Toggl Track
- Notion
- Fluent Migrator
- Yeoman Templating
- PlantUML
- Postman
- MyGet (package management)
- GDPR Compliance