Published

All published posts

2667 posts latest post 2026-09-24 simple view
Publishing rhythm
Sep 2026 | 29 posts

Setting Parameters in kedro

Parameters are a place for you to store variables for your pipeline that can be accessed by any node that needs it, and can be easily changed by changing your environment. Parameters are stored in the repository in yaml files. https://youtu.be/Jj5cQ5bqcjg What is Kedro 👆 Unsure what kedro is? Check out this post. parameters files # You can have multiple parameters files and choose which ones to load by setting your environment. By default kedro will give you a and parameters file. base # The base environment should contain all of the default values you want to run. NOTE base will always be loaded first. accessing parameters # Parameters can be accessed through context or through the catalog. Generally when you are working with nodes it will be loaded through the catalog. Loding with the context. Loading with the catalog. Loading a specific key with the catalog. using parameters in nodes # Here is an example from the complete spaceflights demo. The entire parameters dict is passed in, t…

Writing your first kedro Nodes

https://youtu.be/-gEwU-MrPuA Before we jump in with anything crazy, let’s make some nodes with some vanilla data structures. import node # You will need to import node from kedro.pipeline to start creating nodes. func # The is a callable that will take the and create the. inputs / outputs # Inputs and outputs can be None, a single catalog entry as a string, mutiple catalog entries as a List of strings, or a dictionary of strings where the key is the keyword argument of the func and the value is the catalog entry to use for that keyword. our first node # Sometimes in our pipelines our data is coming from an api where we already have python functions built to pull with. Thats ok, kedro supposrts that with. second node # Now we have some data to work from, lets use that as our input. Multiple Inputs # Kedro can take lists or dicts as either input or output when your function needs more than one input or output. Links # all of my kedro articles kedro playlist on YouTube node docs first_nod…
I’m impressed by skimpy from aeturrell. skimpy is a light weight tool that provides summary statistics about variables in data frames within the console.
I’m really excited about vim-startuptime, an amazing project by dstein64. It’s worth exploring! A plugin for viewing Vim and Neovim startup event timing information.
I like wbthomason’s project packer.nvim. A use-package inspired plugin manager for Neovim. Uses native packages, supports Luarocks dependencies, written in Lua, allows for expressive config
Looking for inspiration? vim-matchup by andymass. vim match-up: even better % 👊 navigate and highlight matching words 👊 modern matchit and matchparen. Supports both vim and neovim + tree-sitter.

Running your Kedro Pipeline from the command line

Running your kedro pipeline from the command line could not be any easier to get started. This is a concept that you may or may not do often depending on your workflow, but its good to have under your belt. I personally do this half the time and run from ipython half the time. In production, I mostly use docker and that is all done with this cli. https://youtu.be/ZmccpLy-OEI What is Kedro 👆 Unsure what kedro is? Check out this post. Kedro run # To run the whole darn project all we need to do is fire up a terminal, activate our environment, and tell kedro to run. Specific Pipelines # Running a sub pipeline that we have created is as easy as telling kedro which one we want to run. Single Nodes # While developing a node or a small list of nodes in a larger pipeline its handy to be able to run them one at a time. Besides the use case of developing a single node I would not reccomend leaning very heavy on running single nodes, let the DAG do the work of figuring out which nodes to run for y…

kedro Virtual Environment

Avoid serious version conflict issues, and use a virtual environment anytime you are running python, here are three ways you can setup a kedro virtual environment. https://youtu.be/ZSxc5VVCBhM conda venv pipenv conda # I prefer to use conda as my virtual environment manager of choice as it give me both the interpreter and the packages I install. I don’t have to rely on the system version of python or another tool to maintain python versions at all, I get everything in one tool. stores environment in a root directory i.e. conda can use its own way to manage environments the python interpreter is packaged with the environment virtualenv # Virtual env (venv) is another very respectable option that is built right into python, and requires no additional installs or using a different distribution of pytyhon. environments are typically stored in the project directory does not package the interpreter pipenv # Pipenv is another virtual enviroment tool that comes with its own system for managing…

Kedro Install

Kedro comes with an command to install and manage all of your projects dependencies. https://youtu.be/IWimEs-hHQg cd into your project directory and activate env # You must start by having your kedro project either cloned down from an existing project or created from kedro new. Then activate your environment. Kedro New this post covers kedro new kedro Virtual Environment This post covers creating your virtual environment for kedro install kedro # Make sure you have kedro installed in your current environment, if you dont already have it. pip-tools # Kedro uses the package under the hood to pin dependencies in a very robust way to ensure that the project will continue to work on everyone’s machine day, including production, day in and day out. No matter what happens to the dependencies you have installed. pip-compile # The command that kedro uses from is. It will look at what you have in a file, compile the dependencies down to exact versions, and create a requirements.txt that is fully…

Kedro Git Init

Immediately after, before you start running or your first line of code the first thing you should always do after getting a new kedro template created is to. https://youtu.be/IGba3ytf_6U git init # Its as simple as these three commands to get started. I don’t care if this project is for learning, if it will never have a remote or not, use git.

Kedro New

https://youtu.be/uqiv5LAiJe0 Kedro new is simply a wrapper around the cookiecutter templating library. The kedro team maintains a ready made template that has everything you need for a kedro project. They also maintain a few kedro starters, which are very similar to the base template. What is Kedro Unsure what kedro is, Check out yesterdays post on What is Kedro. pipx # I reccomend using when running kedro new. is designed for system level cli tools so that you do not need to maintain a virtual environment or worry about version conflicts, manages the environment for you. The kedro team does not reccomend in their docs as they already feel like there is a bit of a tool overload for folks that may be less familiar with I like using as it gives you better control over using a specific version or always the latest version, unlike when you run what you have on your system depends on when you last installed or upgraded. Kedro New # The kedro team also has a set of starters, by passing in yo…

What is Kedro

Kedro is an unopinionated Data Engineering framework that comes with a somewhat opinionated template. It gives the user a way to build pipelines that automatically take care of io through the use of abstract that the user specifies through entries. These entries are loaded, ran through a function, and saved by. The order that these are executed are determined by the, which is a DAG. It’s the ’s job to manage the execution of the. https://youtu.be/Wf4rnFsaFFU What is Kedro This is an updated version of my original what-is-kedro article Hot Take # If you are doing a series of operations to data with python, especially if you are using something as supported as pandas, you should be using a framework that gives you a pipeline as a DAG and abstracts io. Orchestrators # Like I said, is unopinionated it does determine where or how your data should be ran. The kedro team does support the following Orchestrators with very little add on to the base template. Argo Workflows Prefect Kubeflow Work…